Letter analysis method based on language large model and HOG features

Through the letter analysis method based on language big model and HOG characteristics, the problem of inaccurate text analysis of letters and inefficient seal authenticity detection is solved, and the comprehensive understanding of letter content and efficient detection of seal authenticity is achieved, which improves analysis efficiency and accuracy.

CN120354946APending Publication Date: 2025-07-22ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495651.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the inaccurate text analysis of letters and inefficient seal authenticity detection methods are time-consuming and labor-intensive and error-prone, and the technical threshold for forging seals is reduced, making it difficult to accurately identify authenticity.

Method used

A letter analysis method based on language big model and HOG features is adopted. By obtaining letter text and seal data, the language big model and machine learning model are trained, and combined with telegraph templates and HOG feature detection, the letter content understanding and the authenticity of seals are realized.

Benefits of technology

It improves the accuracy of the content of the letter and the verification ability of the seal legality, realizes efficient and accurate letter analysis and seal authenticity detection, and solves the shortcomings in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354946A_ABST
    Figure CN120354946A_ABST
Patent Text Reader

Abstract

The invention discloses a letter analysis method based on a language large model and HOG features, and the method comprises the steps: obtaining a letter text and letter seal data, and obtaining a letter text data set and a true and false seal data set; training a preset language big model and a machine learning model based on the letter text data set and the true and false seal data set, inputting the question of the reviewer and the to-be-detected seal into the trained language big model and machine learning model, and outputting to obtain a language answer and a true and false detection result; and performing fine tuning on the language large model, designing a prompter template, and returning the output of the language large model and the machine learning model to the client. According to the letter text analysis method based on the language large model and the letter seal authenticity detection method based on the HOG feature, the main content of various letters is extracted, and the verification capability of the validity of the seal in the letters is improved; the problems that letter text analysis is inaccurate and a seal authenticity detection method is low in efficiency exist.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of letter analysis, and in particular to a letter analysis method based on a large language model and HOG features. Background Art

[0002] With the continuous advancement of big data technology, artificial intelligence, natural language processing (NLP) technology, and machine learning algorithms, more and more industries are beginning to explore intelligent and automated solutions. However, there are still some technical challenges in the analysis of correspondence texts, such as how to accurately understand the complex meanings in the text, how to efficiently process large amounts of information, and how to ensure that the automated system can achieve the same accuracy as manual processing. Therefore, although there have been some initial attempts at intelligence, overall, the seamless transition from "manual" to "intelligent" has not yet been fully achieved, and most companies in the industry are still seeking appropriate technological breakthroughs and system optimization solutions.

[0003] The current letter text analysis work has not yet achieved the transition from "manual" to "intelligent". It is time-consuming and labor-intensive, prone to errors and omissions, and has problems such as long response time and limitations of physical conditions.

[0004] With the rapid development of science and technology, especially the remarkable progress in photoengraving and computer simulation technology, the technical threshold for counterfeiting seals has been greatly reduced. With the maturity of photoengraving technology, the refined reproduction of images has become simpler and faster, and almost any form of seal pattern and details can be perfectly reproduced, making it almost impossible to distinguish the authenticity with the naked eye.

[0005] At the same time, the rapid development of computer simulation technology has enabled counterfeiters to use advanced image editing software to design, process and modify seals, creating extremely realistic counterfeit seals. The combination of computer-aided design (CAD) and three-dimensional printing technology also allows counterfeiters to quickly produce standard counterfeit seals at a lower cost, and even imitate complex seal textures, fonts and details, greatly improving the accuracy and simulation of counterfeit seals.

[0006] Therefore, building an efficient and accurate method for letter text analysis and seal authenticity detection has become an important issue that needs to be solved urgently.

[0007] In the prior art, there are problems of inaccurate letter text analysis and inefficient seal authenticity detection methods. Summary of the invention

[0008] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a correspondence analysis method based on a language large model and HOG features. Through the correspondence text analysis method based on the language large model and the authenticity detection method of correspondence seals based on HOG features, the main content of various correspondences is refined and the effect of improving the verification ability of the legality of seals in correspondences is achieved, so as to solve the problems of inaccurate correspondence text analysis and inefficient seal authenticity detection methods in the prior art.

[0009] To achieve the above object, the present invention provides the following technical solutions: A correspondence analysis method based on a language large model and HOG features, comprising the following steps: obtaining correspondence text and correspondence seal data, and performing preprocessing to obtain a correspondence text data set and a genuine and fake seal data set; training a preset language large model and a machine learning model based on the correspondence text data set and the genuine and fake seal data set, inputting the questions of reviewers and the seals to be detected into the trained language large model and machine learning model, and outputting a language answer and a authenticity detection result; fine-tuning the language large model, and at the same time designing a prompt template, and returning the outputs of the language large model and the machine learning model to the client.

[0010] In a preferred embodiment, the methods for obtaining the correspondence text data set and the genuine and fake seal data set are specifically as follows: collecting correspondence texts, policy information and industry data in various different formats; preprocessing the correspondences to remove redundant information, and completing format sorting and content rule verification to obtain a correspondence text data set; collecting correspondence seals, and labeling the collected seals as genuine or fake; after unifying the specifications of the labeled seals, a genuine and fake seal data set is obtained.

[0011] In a preferred embodiment, the training of the preset language large model and machine learning model based on the correspondence text data set and the genuine and fake seal data set is specifically as follows: vectorizing the correspondence text data set, performing segmentation processing according to the actual situation of the text, and storing the vectorized data in a local vector database to obtain a vector database; extracting the multi-dimensional HOG features of the genuine and fake seal images in the genuine and fake seal data set, and storing them in a local computer to complete the construction of a feature data set; training the preset language large model and machine learning model in combination with the vector database and the feature data set.

[0012] In a preferred embodiment, the process of inputting the reviewer's questions and the seal to be detected into the trained large language model and machine learning model, and obtaining a language response and authenticity detection result is as follows: vectorize the reviewer's questions, calculate the similarity with the texts in the vector database and match the questions; after obtaining the top K texts with the highest relevance in descending order of similarity, embed the questions and matching texts into a fixed prompt template; extract the HOG features of the seal to be detected and input them into the trained machine learning model to determine the authenticity of the seal.

[0013] In a preferred embodiment, the process of returning the outputs of the large language model and the machine learning model to the client specifically includes: submitting the response content of the large language model in text form and the seal detection result to the human-computer interaction interface for visualization.

[0014] In a preferred embodiment, the process of extracting the multi-dimensional HOG features of genuine and fake seal images in the genuine and fake seal dataset is as follows: grayscale and gamma correct the seal image, and divide the seal image into several cells; calculate the gradient and gradient direction of each pixel in the cell, and statistically generate a gradient histogram; merge several cells into a block, and perform normalization processing on each block to obtain multi-dimensional HOG features.

[0015] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Through a method for analyzing letters based on a large language model and HOG features, specifically a method for analyzing letter texts based on a large language model, with the help of the large language model, a comprehensive understanding and judgment of the letter content can be achieved, and then a list of the main contents of various letters can be refined and formed to assist reviewers in comprehensively controlling and managing the content of supporting documents, effectively solving the problem of inaccurate analysis of letter texts in the prior art.

[0016] 2. Through a method for detecting the authenticity of letter seals based on HOG features, the authenticity of letter seals can be detected efficiently and accurately, thereby improving the ability to verify the legality of seals in letters, and effectively solving the problem of low efficiency of seal authenticity detection methods in the prior art. Brief Description of the Drawings

[0017] Figure 1 It is a flowchart of a method for analyzing letters based on a large language model and HOG features provided by an embodiment of the present application; Figure 2 It is a diagram showing the implementation steps of a method for analyzing letter texts based on a large language model provided by an embodiment of the present application; Figure 3Algorithm flowchart of the method for detecting the authenticity of letter seals based on HOG features provided by the embodiments of the present application; Figure 4 Pre-trained language large model provided by the embodiments of the present application. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "including..." do not exclude the presence of additional identical elements in the process, method, article or device including the said elements.

[0020] Embodiment 1 Figure 1 A method for analyzing letters based on a language large model and HOG features of the present invention is given, including the following steps: Step S1, obtain letter text and letter seal data, and perform preprocessing to obtain a letter text data set and a genuine and fake seal data set; Step S1-1 Obtain the letter text, specifically including: Collect letter texts, policy information and industry data in various different formats; Preprocess the letters to remove redundant information, and complete format sorting and content rule verification to obtain a letter text data set; The method for obtaining the letter seal data is specifically as follows: Collect letter seals, and label the collected seals with true or false labels; After unifying the specifications of the labeled seals, a genuine and fake seal data set is obtained.

[0021] In this embodiment, the letter text formats include PDF, DOCS, TXT, etc.; the seal format is PNG, and the pixel size is 300×300.

[0022] Step S2: Train a pre-set large language model and a machine learning model based on the correspondence text dataset and the genuine and fake seal dataset. Input the reviewer's questions and the seal to be detected into the trained large language model and machine learning model, and output a language answer and a genuine / fake detection result. Step S21: Train a pre-set large language model and a machine learning model based on the correspondence text dataset and the genuine and fake seal dataset to obtain a vector database and a feature dataset, specifically including: Step S21-1: Vectorize the correspondence text dataset, perform segmentation processing according to the actual situation of the text, and store the vectorized data in a local vector database to obtain the vector database. Perform word segmentation on the correspondence text dataset, calculate word vectors for each segmented document using the BERT model, and calculate the TF-IDF value of each word based on the word segmentation results. Combine the word vectors with the TF-IDF values of each word to obtain word vectors with weight information, and obtain the vector database. The IDF calculation formula for word t is as follows:

[0023] Where: M is the total number of training texts; is the number of documents in the training text set that contain word t.

[0024] The calculation formula for TF-IDF is as follows:

[0025] Where: is the word frequency of word t in the i-th text; is the normalization factor.

[0026] Step S21-2: Extract the multi-dimensional HOG features of the genuine and fake seal images in the genuine and fake seal dataset, store them in a local computer, and complete the construction of the feature dataset. Grayscale and gamma correct the seal image, and divide the seal image into several cells; Calculate the gradient and gradient direction of each pixel in the cell, and statistically generate a gradient histogram; Merge several cells into a block, and perform normalization processing on each block to obtain multi-dimensional HOG features.

[0027] In this embodiment, according to the sensitivity of the human eye to the three colors R, G, and B, the weighted average method is used to grayscale the image. Each cell contains 8×8 pixels, and each block is obtained in the form of a sliding window. Each block contains 2×2 cells, and the step size of the sliding window is 1 cell. The L2 norm is used to normalize the block. The involved formulas are as follows: Weighted average method:

[0028] In the formula: D represents the grayscale value after conversion of the pixel point (x, y), and R, G, and B are the components of the three channels of this pixel point.

[0029] Gamma correction:

[0030] In the formula: represents the input grayscale value, represents the output grayscale value, and gamma takes 1 / 2.2.

[0031] L2 norm:

[0032] In the formula: v represents the histogram within the block, is a small constant used to avoid the denominator being zero.

[0033] Step S21-3: Train the preset large language model and machine learning model by combining the vector database and the feature dataset.

[0034] Vectorization refers to converting words, sentences, or paragraphs in text into digital vectors. Through vectorization, the information in the text can be processed by digital models. In this step, the text content will be converted into a high-dimensional numerical representation. For example, each word is converted into a "vector".

[0035] Segmentation processing is to make each part of the text more controllable and easy to process. For example, if a letter is particularly long, it can be divided into several paragraphs, sentences, or even smaller units. This is to avoid the input text being too long and difficult to process in the model, and segmentation also helps to maintain context information.

[0036] The goal of word segmentation is to break down sentences or paragraphs into meaningful basic units.

[0037] Step S22: Input the questions of the reviewers and the seals to be detected into the trained large language model and machine learning model, and output the language answers and authenticity detection results, specifically including: Step S22-1: Vectorize the questions of the reviewers and calculate the similarity with the texts in the vector database for question matching. When calculating the similarity, avoid directly performing topic mapping on short texts. Instead, calculate the probability of generating the short text based on the topic distribution of the long text as their similarity.

[0038] The calculation formula for similarity is as follows:

[0039] In the formula: q represents Query, c represents content, w represents the words in q, represents the k-th topic.

[0040] Step S22-2: After obtaining the top K texts with the highest relevance in descending order of similarity, embed the question and the matching text into a fixed prompt template. Step S22-3: Extract the HOG features of the seal to be detected and input them into the trained machine learning model to determine the authenticity of the seal.

[0041] In this embodiment, the model includes: LLM (Large Language Model): These models take text strings as input and return text strings as output. They are the backbone of many language model applications.

[0042] Chat Model: The chat model is supported by the large language model but has a more structured API. They take a list of chat messages as input and return chat messages. This makes it easy to manage the conversation history and maintain context.

[0043] Text Embedding Models: These models take the correspondence text as input and return a floating-point list representing the text embedding. These embeddings can be used for tasks such as document retrieval, clustering, and similarity comparison.

[0044] Support Vector Machine (SVM) is a machine learning algorithm used for classification and regression analysis. The core idea of SVM is to find a hyperplane that can separate the sample points of different classes as much as possible and maximize the margin between the two classes. Therefore, it can well distinguish genuine and fake seals.

[0045] Step S3: Fine-tune the large language model, and at the same time design a prompt template, and return the outputs of the large language model and the machine learning model to the client.

[0046] Step S31: Fine-tune the large language model, specifically including: Perform implicit low-rank transformation on the weight matrix of the large model, specifically including: Assume that W represents the weight matrix in the neural network layer. Using common backpropagation, the weight update can be obtained , which is obtained by multiplying the negative gradient of the loss by the learning rate. The calculation formula is as follows:

[0047] The updated weight is expressed as:

[0048] In the formula: is the updated weight, is the updated weight, is the weight before update.

[0049] In this embodiment, Lord is selected to fine-tune the language large model, that is, perform implicit low-rank transformation on the weight matrix of the large model. A bypass structure is added to the network, and the bypass is the multiplication of two matrices A and B. The dimension of matrix A is , and the dimension of matrix B is , where , and generally r takes 1, 2, 4, 8. Then the number of parameters of this bypass will be much smaller than the parameter W of the original network. During LoRA training, the parameter W of the original network is frozen, and only the bypass parameters A and B are trained.

[0050] Step S32, design a prompt template according to the query purpose, specifically including: Based on the principle of In Context Learning (ICL), the prompt template can be expressed as the following formula:

[0051] In the formula: represents the prompt instruction for the specific task, represents the text from which information is to be obtained.

[0052] Step S33, return the outputs of the language large model and the machine learning model to the client, specifically including: Submit the response content of the language large model in text form and the seal detection result to the human-computer interaction interface for visualization.

[0053] In this embodiment, the execution result obtained in S3 is combined with the seal detection result obtained in S2 to obtain a new response statement, and it is returned to the graphical interface of the user end.

[0054] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0055] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0056] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0057] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0058] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0059] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A correspondence analysis method based on a large language model and HOG features, characterized in that, It includes the following steps: Obtain the correspondence text and correspondence seal data, and perform preprocessing to obtain the correspondence text dataset and the genuine and fake seal dataset; Based on the correspondence text dataset and the genuine and fake seal dataset, train the preset language large model and machine learning model, input the questions of the reviewers and the seals to be detected into the trained language large model and machine learning model, and output the language answers and authenticity detection results; Fine-tune the language large model, and at the same time design a prompter template, and return the outputs of the language large model and machine learning model to the client.

2. The method for analyzing letters based on a large language model and HOG features according to claim 1, wherein For the correspondence text dataset and the genuine and fake seal dataset, the specific acquisition method is as follows: Collect correspondence texts, policy information, and industry data in various different formats; preprocess the correspondence to remove redundant information, and complete format sorting and content rule verification to obtain the correspondence text dataset; Collect correspondence seals, and label the collected seals as genuine or fake; After unifying the specifications of the labeled seals, obtain the genuine and fake seal dataset.

3. The method for analyzing letters based on a large language model and HOG features according to claim 2, wherein The training of the preset language large model and machine learning model based on the correspondence text dataset and the genuine and fake seal dataset is specifically as follows: Vectorize the correspondence text dataset, perform segmentation processing according to the actual situation of the text, and store the vectorized data in the local vector database to obtain the vector database; Extract the multi-dimensional HOG features of the genuine and fake seal images in the genuine and fake seal dataset, store them in the local computer, and complete the construction of the feature dataset; Combine the vector database and the feature dataset to train the preset language large model and machine learning model.

4. The method for analyzing letters based on a large language model and HOG features according to claim 3, characterized in that, The input of the questions of the reviewers and the seals to be detected into the trained language large model and machine learning model, and the output of the language answers and authenticity detection results is specifically as follows: Vectorize the questions of the reviewers, and calculate the similarity and question matching with the texts in the vector database; After obtaining the top K texts with the highest correlation in descending order of similarity, embed the questions and matching texts into a fixed prompt template; Extract the HOG features of the seal to be detected and input them into the trained machine learning model to judge the authenticity of the seal.

5. The method for analyzing letters based on a large language model and HOG features according to claim 4, wherein, The return of the outputs of the language large model and machine learning model to the client specifically includes: Submit the reply content of the language large model in text form and the seal detection result to the human-computer interaction interface for visualization.

6. The method for analyzing letters based on a large language model and HOG features according to claim 5, wherein The extraction of the multi-dimensional HOG features of the genuine and fake seal images in the genuine and fake seal dataset is specifically as follows: Grayscale and gamma correct the seal image, and divide the seal image into several cells; Calculate the gradient and gradient direction of each pixel in the cell, and statistically generate a gradient histogram; Merge several cells into a block, and perform normalization processing on each block to obtain multi-dimensional HOG features.

7. The method for analyzing letters based on a large language model and HOG features according to claim 6, wherein, The fine-tuning of the language large model is specifically as follows: Perform implicit low-rank transformation on the weight matrix of the large model, specifically including: Preset W to represent the weight matrix in the neural network layer, and obtain the updated weight by multiplying the negative gradient of the loss by the learning rate. The calculation formula is as follows: The updated weight is expressed as: Wherein: is the updated weight, is the updated weight, is the weight before update.

8. The method for analyzing letters based on a large language model and HOG features according to claim 7, wherein, The design of the prompter template is specifically as follows: Design a prompt template according to the query purpose. The prompt template can be expressed by the following formula: Wherein: represents a prompt instruction for a specific task, represents the text from which information is to be obtained.

9. The method for analyzing letters based on a large language model and HOG features according to claim 8, wherein The method for obtaining the vector database is specifically as follows: Perform word segmentation on the correspondence text dataset, calculate word vectors for each document after word segmentation using the BERT model, and calculate the TF-IDF value of each word based on the word segmentation results. Combine the word vectors with the TF-IDF values of each word to obtain word vectors with weight information, and obtain the vector database; The IDF value of each word is calculated according to the following formula: Where: M is the total number of training texts; is the number of documents in the training text set in which the word t appears; The calculation formula of TF-IDF is as follows: In the formula: is the word frequency of word t in the i-th text; is the normalization factor.

10. The method for analyzing letters based on a large language model and HOG features according to claim 9, characterized in that, The specific method for vectorizing the questions of reviewers and calculating the similarity with the text in the vector database is as follows: Where: q represents Query, c represents content, w represents the word in q, represents the k-th topic.