Power enterprise knowledge base updating and maintaining system and method based on NLP large model
Through the NLP large model of the power enterprise knowledge base update and maintenance system, the problem of inaccurate extraction of handwritten information in the power enterprise knowledge base is solved, the accurate extraction and entity recognition of handwritten information are achieved, and the efficiency of updating the knowledge base of power enterprise is improved.
Patent Information
- Application Number
- CN202510521625.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the problem of inaccurate extraction of handwritten information during the update process of the power enterprise knowledge base.
The power enterprise knowledge base update and maintenance system based on the NLP large model is adopted, including the knowledge picture segmentation module, the character integrity judgment module, the general library comparison module and the power word library comparison module. By segmenting the regional knowledge pictures, the integrity and matching degree of character slices are judged, the word set to be compared is generated and the matching degree is analyzed, and whether it is manually reviewed.
It improves the accuracy of the extraction of handwritten information in the knowledge base of power enterprises, ensures the accuracy of character recognition and entity recognition of handwritten documents, and improves the efficiency of the update and maintenance of the knowledge base of power enterprises by NLP.
Smart Images

Figure CN120388384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a power enterprise knowledge base update and maintenance system and method based on an NLP large model. Background Art
[0002] With the continuous update of a large number of technical documents, operation data, and policies and regulations in power enterprises, the traditional manual maintenance method is inefficient and error-prone. By introducing an NLP (Natural Language Processing) large model, the system can automatically process and understand massive text data, perform knowledge extraction, classification, and update, and maintain the accuracy and timeliness of the knowledge base in real time. This method not only improves work efficiency but also provides a more intelligent solution for fields such as decision-making support and fault diagnosis, promoting the digital transformation of power enterprises and the development of smart energy.
[0003] Existing methods rely on manual input and editing, suffering from problems such as low efficiency, high cost, and easy errors. With the development of artificial intelligence and natural language processing technologies, some power enterprises have attempted to use NLP technology for automated text processing and knowledge extraction. Current technical means such as information extraction (IE, Information Extraction), named entity recognition (NER, Named Entity Recognition), and text classification can, to a certain extent, achieve automated update of literature and data, but still face problems such as insufficient understanding of the professional field, poor model adaptability, and information fusion issues. Therefore, the solution based on the NLP large model can further improve accuracy, intelligence level, and real-time update ability.
[0004] For example, the natural language processing method, device, equipment, and storage medium disclosed in the invention patent announcement with the publication number: CN114428788B includes: performing entity recognition on a natural language description to obtain an entity string sequence, where each entity string in the entity string sequence is a label of an entity in the natural language description in the target language; aggregating the entity string sequence based on a target grammar, and performing aggregation fallback on the entity string sequence that fails to aggregate, where the target grammar is the grammar of the target language; performing re-aggregation based on the aggregation fallback result to obtain semantic information to be parsed, where the semantic information to be parsed is the aggregation relationship between entity strings in the entity character sequence; and parsing the semantic information to be parsed to obtain a target language description.
[0005] For example, a knowledge enhancement method and system for a large language model announced in the invention patent announcement with the announcement number of CN118779470B includes: obtaining vertical domain data information to be processed; constructing a vertical domain knowledge base based on the vertical domain data information to be processed; extracting topic words in the vertical domain knowledge base based on the topic recognition method, and constructing a Monte Carlo tree structure corresponding to the vertical domain knowledge base; identifying topic words for the Monte Carlo tree structure in the vertical domain based on the topic word identification method to determine the topic word categories of each text block in the vertical domain knowledge base; processing the topic word categories and a preset user question to obtain a sentence sorting table corresponding to the text block; generating a prompt word template based on the sentence sorting table and the user question, and constructing an enhanced large language model based on the prompt word template.
[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, it is found that the above technology has at least the following technical problems:
[0007] In the prior art, during the update process of uploading power enterprise knowledge documents, the character extraction of the handwritten documents scanned in the power enterprise knowledge base may be inaccurate, resulting in inaccurate extraction of the handwritten information that needs to be updated in the power enterprise knowledge base. Summary of the Invention
[0008] The embodiments of the present application provide a power enterprise knowledge base update and maintenance system and method based on an NLP large model, which solve the problem of inaccurate extraction of handwritten information that needs to be updated in the power enterprise knowledge base in the prior art, and achieve an improvement in the accuracy of extracting handwritten information that needs to be updated in the power enterprise knowledge base.
[0009] The embodiments of the present application provide a power enterprise knowledge base update and maintenance system based on an NLP large model, including: a knowledge picture segmentation module, a character integrity judgment module, a general library comparison module, and a power word library comparison module: Among them, the knowledge picture segmentation module is used to perform picture processing on the regional knowledge picture to obtain corresponding character segmentation data, and perform segmentation processing on the regional knowledge picture based on the character segmentation data to obtain character slices, and the character slices include single character slices and non-single character slices; the character integrity judgment module is used to judge whether to perform manual verification through the proportion of single character slices. If no manual verification is performed, the character morphology of each character slice is evaluated to obtain character morphology data, and based on this, the integrity of each character slice is judged; the general library comparison module is used to compare the character slices with the general character library to obtain a set of characters to be compared and character recognition data, and perform character recognition analysis according to the character recognition data to obtain a set of words to be compared; the power word library comparison module is used to compare the set of words to be compared with the power enterprise knowledge word library, and perform a matching degree analysis according to the word library data obtained after the comparison to judge whether to perform manual review.
[0010] Further, the character segmentation data includes the horizontal character spacing, the vertical character spacing, and the character edge density; the character form data includes the character confidence and the character connectivity; the character recognition data includes the character matching rate, the character overlap rate, the character matching completion time, and the mis-matching rate; the character overlap rate represents the success rate of the match by analyzing the overlapping part of the shape of the character and the characters in the library; the thesaurus data includes the context matching degree, the thesaurus matching success rate, the thesaurus matching completion time, and the thesaurus word null value rate; the context matching degree represents the matching degree of the character with the context of its preceding and following characters; the thesaurus word null value rate represents the proportion of the words not matched in the thesaurus in the total vocabulary.
[0011] Further, the specific process of segmenting the regional knowledge picture into character slices based on the character segmentation data is as follows: perform data normalization processing on the character segmentation data obtained by performing picture processing on the regional knowledge picture, and obtain the character segmentation weight factors and the minimum reference ratio from a preset database. The character segmentation weight factors include the horizontal character weight factor, the vertical character weight factor, and the character edge weight factor; obtain the segmentation difficulty quantization coefficient based on the horizontal character spacing, the vertical character spacing, and the minimum reference ratio. The segmentation difficulty quantization coefficient represents the quantization data of the combined influence of the horizontal character spacing and the vertical character spacing on the difficulty of segmenting the handwritten document; perform preliminary recognition on the regional knowledge picture to obtain preliminary recognized characters and number them, and obtain the character segmentation data of each preliminary recognized character; perform weighted processing on the character segmentation data in combination with the character segmentation weight factors and then perform character differentiation processing to obtain the character differentiation processing result, and perform arithmetic processing on the character differentiation processing result and the segmentation difficulty quantization coefficient to obtain the character differentiation coefficient. The character differentiation processing is used to establish the data mapping relationship between the character segmentation data and the character differentiation coefficient. The character differentiation coefficient represents the quantization data of the combined influence of the horizontal character spacing, the vertical character spacing, and the character edge density on the independence degree between characters; if the character differentiation coefficient is not less than the character differentiation threshold in the preset database, perform special-shaped column segmentation on the regional knowledge picture to obtain the corresponding column character slices, otherwise obtain the character segmentation data of the next horizontally adjacent preliminary recognized character until the corresponding character differentiation coefficient is not less than the character differentiation threshold for special-shaped column segmentation; perform row segmentation based on the character differentiation coefficients of each preliminary recognized character in each column character slice to obtain character slices.
[0012] Further, the specific steps for determining whether to perform manual verification based on the single-character slice ratio are as follows: A1. Compare the single-character slice ratio with a preset single-character slice ratio. If the single-character slice ratio is less than the preset single-character slice ratio, perform manual verification; otherwise, execute A2. Manual verification means feeding the corresponding character slice to a preset verification personnel to obtain the character of the corresponding character slice. A2. Perform character form evaluation on each character slice to obtain character form data. If the character slice does not have character form data, delete the corresponding character slice; otherwise, execute A3. The character form evaluation is used to judge the integrity of the character according to the form of the character. A3. If the character form data of each character slice is not less than the corresponding reference form data, no manual verification is performed; otherwise, combine the character slice with the horizontally adjacent character slice to obtain a non-single character slice. The reference form data includes the maximum character confidence and the maximum character connectivity.
[0013] Further, the specific process of performing character recognition analysis based on the character recognition data to obtain a word set to be compared is as follows: Number the character slices and compare them with a general character library to obtain a character set; Perform irrelevant part-of-speech recognition on the character set and remove the irrelevant words in the character slices to obtain a word set to be compared and character recognition data, and obtain a character recognition coefficient according to the character recognition data and the word correction factor; If the character recognition coefficient is not less than the character recognition threshold in the preset database, combine the character slices according to the character slice numbers to obtain the corresponding word set to be compared; otherwise, optimize the character discrimination threshold; If the character recognition coefficient corresponding to the character slice obtained based on the optimized character discrimination threshold is less than the character recognition threshold, perform manual verification; otherwise, obtain the corresponding word set to be compared.
[0014] Further, the specific process for obtaining the character recognition coefficient is as follows: Obtain a character recognition weight factor and a maximum matching time from a preset database. The character recognition weight factor includes a character matching weight factor, a character overlap weight factor, a character matching time weight factor, and a mis-matching rate weight factor; Obtain a word correction difference according to the word correction factor; Perform character recognition processing on each character recognition data respectively and then perform weighted processing in combination with the character matching weight factor to obtain a character recognition processing result. The character recognition processing is used to unify the logical relationship between the character recognition data and the character recognition coefficient; Multiply the character recognition processing result by the word correction difference to obtain the character recognition coefficient. The character recognition coefficient represents the quantitative data of the accuracy degree of character recognition after character segmentation jointly by the character matching rate, the character overlap rate, the character matching completion time, and the mis-matching rate.
[0015] Further, the specific process for obtaining the word correction factor is as follows: Obtain the correction data and the correction weights in the preset database. The correction data includes morphological similarity and ambiguity, and the correction weights include morphological weight and ambiguity weight. Input the correction data and the correction weights into the correction model to obtain the corresponding word correction factor. The word correction factor represents the quantitative data of the combined effect of morphological similarity and ambiguity on reducing the matching degree of semantic comparison pairs. The correction model is used to establish the logical relationship between the correction data and the word correction factor.
[0016] Further, the process of comparing the word set to be compared with the power enterprise knowledge word bank and analyzing the matching degree according to the word bank data obtained after the comparison is as follows: Obtain the word bank matching weight factor and the maximum completion time from the preset database. The word bank matching weight factor includes the context matching degree weight factor, the word bank matching weight factor, the word bank completion time weight factor, and the word bank word null value weight factor. After performing word bank comparison processing on each word bank data respectively, perform weighted processing in combination with the word bank matching weight factor to obtain the word bank comparison processing result. The word bank comparison processing is used to establish the logical relationship between the word bank data and the power enterprise word bank matching coefficient. Perform arithmetic processing on the word bank comparison processing result and the word correction difference to obtain the power enterprise word bank matching coefficient. The power enterprise word bank matching coefficient represents the quantitative data of the combined effect of context matching degree, word bank matching success rate, word bank matching completion time, and word bank word null value rate on the matching degree between the recognized characters and the vocabulary in the power enterprise professional knowledge word bank.
[0017] Further, the specific process for determining whether to perform manual review is as follows: Compare the power enterprise word bank matching coefficient with the power word bank threshold obtained from the preset database. If the power enterprise word bank matching coefficient is not less than the power word bank threshold, combine all character slices in the order of their numbers according to the comparison result of the word set to be compared, and at the same time automatically generate a text report based on the NLP large model for entity recognition. If the power enterprise word bank matching coefficient is less than the power word bank threshold, feedback the character slices in the word set to be compared without comparison results to the preset staff for manual verification.
[0018] In the embodiments of the present application, a method for updating and maintaining a knowledge base of power enterprises based on an NLP large model is provided, including the following steps: performing image processing on regional knowledge images to obtain corresponding character segmentation data, and performing segmentation processing on the regional knowledge images based on the character segmentation data to obtain character slices, where the character slices include single-character slices and non-single-character slices; judging whether to perform manual verification based on the proportion of single-character slices, and if not, performing character form evaluation on each character slice to obtain character form data, and accordingly judging the integrity of each character slice; comparing the character slices with a general character library to obtain a set of characters to be compared and character recognition data, and performing character recognition analysis based on the character recognition data to obtain a set of words to be compared; comparing the set of words to be compared with the knowledge word library of power enterprises, and performing matching degree analysis based on the word library data obtained after comparison to judge whether to perform manual review.
[0019] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0020] 1. By performing segmentation processing on regional knowledge images to obtain character slices, then judging whether to perform manual verification based on the proportion of single-character slices, then comparing the character slices with a general character library to obtain a set of characters to be compared and performing character recognition analysis to obtain a set of words to be compared, and finally comparing the set of words to be compared with the knowledge word library of power enterprises to perform matching degree analysis to judge whether to perform manual review, more accurate handwritten information extraction is obtained, thereby improving the accuracy of handwritten information extraction required for updating the knowledge base of power enterprises, and effectively solving the problem of inaccurate extraction of handwritten information required for updating the knowledge base of power enterprises in the prior art.
[0021] 2. By numbering the character slices and comparing them with a general character library to obtain a character set, then removing irrelevant words in the character slices to obtain a set of characters to be compared and character recognition data, obtaining a character recognition coefficient according to the character recognition data and a word correction factor, if the character recognition coefficient is not less than the character recognition threshold in the preset database, combining the character slices in the set of characters to be compared according to the character slice numbers to obtain a corresponding set of words to be compared, otherwise optimizing the character discrimination threshold and judging whether to perform manual verification, thereby further evaluating the accuracy of character recognition, and thus obtaining a more accurate set of words to be compared.
[0022] 3. By obtaining a word library matching weight factor and a maximum completion time, performing word library comparison processing on each word library data respectively and then performing weighted processing in combination with the word library matching weight factor to obtain a word library comparison processing result, and then performing a product operation on the word library comparison processing result and a word correction difference to obtain a power enterprise word library matching coefficient, thereby more accurately evaluating the degree of conformity between the extracted handwritten characters and the knowledge of power enterprises, and thus improving the semantic understanding ability of the NLP large model for the extracted handwritten character text. Brief Description of the Drawings
[0023] Figure 1 This is a schematic structural diagram of the power enterprise knowledge base update and maintenance system based on the NLP large model provided by the embodiment of the present application. Detailed Embodiment
[0024] In the embodiment of the present application, by providing a power enterprise knowledge base update and maintenance system and method based on the NLP large model, the problem of inaccurate extraction of handwritten information that needs to be updated in the existing power enterprise knowledge base is solved. By performing image processing on the regional knowledge image to obtain character segmentation data, and based on the character segmentation data, performing segmentation processing on the regional knowledge image to obtain character slices. Then, by numbering the character slices and comparing them with the general character library to obtain a character set, and then removing irrelevant words in the character slices to obtain a character set to be compared and character recognition data. According to the character recognition data and the word correction factor, a character recognition coefficient is obtained. If the character recognition coefficient is not less than the character recognition threshold in the preset database, the character slices in the character set to be compared are combined according to the character slice numbers to obtain the corresponding word set to be compared. Otherwise, the character discrimination threshold is optimized and it is judged whether manual verification is required. Finally, the word set to be compared is compared with the power enterprise knowledge word library to obtain word library data, and a matching degree analysis is performed to judge whether manual review is required, realizing the improvement of the accuracy of extracting handwritten information that needs to be updated in the power enterprise knowledge base.
[0025] The technical solution in the embodiment of the present application is to solve the above problem of inaccurate extraction of handwritten information that needs to be updated in the power enterprise knowledge base. The general idea is as follows:
[0026] By performing segmentation processing on the regional knowledge image to obtain character slices, then judging whether manual verification is required through the single character slice ratio, then comparing the character slices with the general character library to obtain a character set to be compared and performing character recognition analysis to obtain a word set to be compared, and finally comparing the word set to be compared with the power enterprise knowledge word library to perform a matching degree analysis to judge whether manual review is required, achieving the effect of improving the accuracy of extracting handwritten information that needs to be updated in the power enterprise knowledge base.
[0027] In order to better understand the above technical solution, the above technical solution will be described in detail below in combination with the accompanying drawings of the specification and specific embodiments.
[0028] As Figure 1As shown in the figure, it is a schematic structural diagram of the power enterprise knowledge base update and maintenance system based on the NLP large model provided by the embodiments of the present application. The power enterprise knowledge base update and maintenance system based on the NLP large model provided by the embodiments of the present application includes: a knowledge picture segmentation module, a character integrity judgment module, a general library comparison module, and a power word library comparison module. Among them, the knowledge picture segmentation module is used to perform picture processing on the regional knowledge picture to obtain corresponding character segmentation data, and perform segmentation processing on the regional knowledge picture based on the character segmentation data to obtain character slices, where the character slices include single-character slices and non-single-character slices. The character integrity judgment module is used to judge whether to perform manual verification through the proportion of single-character slices. If no manual verification is performed, the character morphology of each character slice is evaluated to obtain character morphology data, and based on this, the integrity of each character slice is judged. The general library comparison module is used to compare the character slices with the general character library to obtain a set of characters to be compared and character recognition data, and perform character recognition analysis based on the character recognition data to obtain a set of words to be compared. The power word library comparison module is used to compare the set of words to be compared with the power enterprise knowledge word library, and perform a matching degree analysis based on the word library data obtained after the comparison to judge whether to perform manual review.
[0029] In this embodiment, through the training of deep learning technology and a large-scale corpus, the NLP large model can significantly improve the computer's ability to understand human language. The large model can identify entities, relationships, and emotions in sentences, so as to achieve a comprehensive understanding of the text meaning. When updating the power enterprise knowledge base and uploading a handwritten document that needs to be updated, the character extraction from the handwritten document may be inaccurate. Through the collaborative action of each module, the character extraction of the handwritten document for updating and uploading the power enterprise knowledge base is realized, and at the same time, the accuracy of the character extraction from the handwritten document is ensured. Moreover, the characters of the handwritten document extracted based on the embodiments of the present application are more in line with the power enterprise knowledge base, improving the accuracy of entity recognition in the characters of the extracted handwritten document by the NLP large model, and further improving the efficiency of the NLP large model in updating and maintaining the power enterprise knowledge base. Among them, the picture processing NLP large model is realized by combining a vision-language model (such as CLIP, Contrastive Language-Image Pretraining), and the character morphology evaluation is evaluated through morphological analysis.
[0030] Specifically, morphological analysis includes dilation and erosion, opening and closing operations, and skeleton extraction. Dilation and erosion help identify the contours, structures, and gaps between strokes of characters. For example, the erosion operation can remove small noise points in characters, while the dilation operation can enhance the strokes of characters. Opening and closing operations are usually used to remove adjacent connected parts of characters and help evaluate whether the characters are complete or have interruptions. Skeleton extraction can extract the skeleton of characters through skeletonization methods (such as thinning algorithms), that is, removing the peripheral parts of characters and only retaining the core structural features, which helps evaluate the basic shape and proportion of characters.
[0031] Furthermore, character segmentation data includes character horizontal spacing, character vertical spacing, and character edge density; character morphology data includes character confidence and character connectivity; character recognition data includes character matching rate, character overlap rate, character matching completion time, and mis - matching rate; the character overlap rate represents the success rate of matching by analyzing the overlapping parts of the shapes of characters and characters in the library; lexicon data includes context matching degree, lexicon matching success rate, lexicon matching completion time, and lexicon word null value rate; the context matching degree represents the degree of matching between a character and the context of its preceding and following characters; the lexicon word null value rate represents the proportion of words not matched in the lexicon among the total vocabulary.
[0032] In this embodiment, the character horizontal spacing represents the distance between a preliminarily recognized character and the next preliminarily recognized character adjacent horizontally, the character vertical spacing represents the distance between a preliminarily recognized character and the next preliminarily recognized character adjacent vertically, the character edge density is obtained by using the Canny algorithm to detect the edges of characters in the image and calculating the ratio of the number of edge pixels to the total number of pixels in the image, and the character horizontal spacing and character vertical spacing are obtained by using contour detection technology to detect the boundaries of each preliminarily recognized character and measuring the spacing between the boundaries of the preliminarily recognized characters.
[0033] It should be added that the character confidence is obtained through modern OCR tools (such as Tesseract, EasyOCR, etc.). The character confidence is usually a floating value between 0 and 1, representing the accuracy of the recognition of the preliminarily recognized character. Generally, the higher the character confidence, the greater the probability that the preliminarily recognized character is correctly recognized; the character connectivity is detected by edge detection (such as Canny edge detection) and connected component analysis (such as findContours in OpenCV) to detect the connectivity of each preliminarily recognized character in the image, and then connected region detection is performed, and the character connectivity of the preliminarily recognized character is calculated by using connected component analysis.
[0034] It should be added that the character matching rate represents the ratio of the number of character segments that match successfully in the general character library to the total number of character segments; the character overlap rate represents the ratio of the number of character segments with overlapping character segment areas among the character segments that fail to match to the total number of character segments, the character matching completion time represents the time from the start of matching the first character segment to the completion of matching all character segments, and the mis-matching rate represents the ratio of the number of character segments that are mis-matched to the total number of character segments. If a character segment matches multiple comparison results during the comparison process, it is regarded as a mis-match.
[0035] It should be added that the context matching degree is judged by comparing the number of character segments that match successfully with the actual applied language models (such as BERT, RoBERTa, T5, etc.). Among them, the ratio of the number of character segments that do not conform to the actual applied language model to the total number of character segments is the context matching degree; the success rate of the thesaurus matching represents the ratio of the number of character segments that match successfully to the total number of character segments; the thesaurus matching completion time represents the time from the start of comparing the first combined character segment to the completion of comparing all combined character segments, and the thesaurus word null value rate represents the ratio of the number of words of the character segment combinations that are not matched in the power enterprise knowledge thesaurus to the number of words of all character segment combinations.
[0036] Through the acquisition of the above data, it provides a data basis for the analysis of character extraction of the handwritten documents updated and uploaded to the power enterprise knowledge base, which is conducive to improving the accuracy of character recognition and extraction of the handwritten documents that need to be updated.
[0037] Furthermore, the specific process of segmenting the regional knowledge picture into character segments based on the character segmentation data is as follows: perform data normalization processing on the character segmentation data obtained by performing picture processing on the regional knowledge picture. The data normalization processing is used to remove the unit and unify the dimension of the character segmentation data. Obtain the character segmentation weight factors and the minimum reference ratio from the preset database. The character segmentation weight factors include the character horizontal weight factor, the character vertical weight factor, and the character edge weight factor. The picture processing includes preprocessing, text region detection, and region segmentation; obtain the segmentation difficulty quantization coefficient based on the character horizontal spacing, the character vertical spacing, and the minimum reference ratio (i.e., ) The segmentation difficulty quantization coefficient represents the quantization data of the combined effect of the horizontal character spacing and the vertical character spacing on the difficulty of segmenting a handwritten document. The regional knowledge picture is initially recognized to obtain initially recognized characters and numbered, and the character segmentation data of each initially recognized character is obtained. After weighting the character segmentation data with the character segmentation weight factor, character differentiation processing is performed to obtain the character differentiation processing result. The character differentiation processing result is operated with the segmentation difficulty quantization coefficient to obtain the character differentiation coefficient. The character differentiation processing is used to establish a data mapping relationship between the character segmentation data and the character differentiation coefficient. The character differentiation coefficient represents the quantization data of the combined effect of the horizontal character spacing, the vertical character spacing, and the character edge density on the independence degree between characters. If the character differentiation coefficient is not less than the character differentiation threshold in the preset database, the regional knowledge picture is segmented into special-shaped columns to obtain the corresponding column character slices. Otherwise, the character segmentation data of the next adjacent initially recognized character horizontally is obtained until the corresponding character differentiation coefficient is not less than the character differentiation threshold for special-shaped column segmentation. Row segmentation is performed based on the character differentiation coefficients of each initially recognized character in each column character slice to obtain character slices.
[0038] The specific limiting expression of the character differentiation coefficient is as follows:
[0039]
[0040] In the formula, n represents the number of the initially recognized character, n = 1, 2,..., N, N represents the total number of initially recognized characters, HC n represents the horizontal character spacing of the nth initially recognized character, VC n represents the vertical character spacing of the nth initially recognized character, RR represents the minimum reference ratio, CED n represents the character edge density of the nth initially recognized character, θ a represents the horizontal character weight factor, θ b represents the vertical character weight factor, θ c represents the character edge weight factor, CDC n represents the character differentiation coefficient of the nth initially recognized character.
[0041] In this embodiment, the preprocessing includes grayscale conversion, denoising, binarization, and skew correction; text region detection means detecting the text part in the image of the knowledge to be extracted, removing the non-text part to obtain the image of the knowledge to be segmented, and obtaining the regional knowledge image after regionally segmenting the image of the knowledge to be segmented. The image of the knowledge to be extracted represents the image formed by scanning a handwritten document, and regional segmentation means segmenting according to the text language type (such as Chinese, English, etc.); irregular column segmentation means segmenting the initially recognized characters in the first vertical column of the text in the regional knowledge image, and the segmentation can be irregularly segmented according to the actual situation of the initially recognized characters to obtain column character slices; the minimum reference ratio is obtained in advance by a preset staff from a preset database, and the minimum reference ratio represents the character spacing ratio with the lowest convenience for segmenting the regional knowledge image; a single-character slice means a character slice in which only one initially recognized character is segmented, and a non-single-character slice means a character slice in which more than one initially recognized character is segmented. Through the method provided in the embodiments of the present application, it helps to more accurately and finely segment the handwritten document, which is beneficial to the accurate extraction of the characters in the handwritten document, and further ensures the reliability of the information extraction of the handwritten documents of power enterprises, thereby improving the efficiency of the NLP large model in updating and maintaining the knowledge base of power enterprises.
[0042] Specifically, the character discrimination coefficient algorithm provided in this embodiment comprehensively analyzes the character segmentation data, the corresponding character segmentation weight factor, and the minimum reference ratio to obtain the character discrimination coefficient. In the formula, as the character horizontal spacing, character vertical spacing, and character edge density (the value range is from 0 to 1, excluding 0 and 1) increase, the corresponding character discrimination coefficient also increases, indicating that the possibility that the corresponding initially recognized character is a single character is higher. At the same time, when the segmentation difficulty quantization coefficient (i.e., ) obtained based on the character horizontal spacing, character vertical spacing, and the minimum reference ratio is larger, it means that the segmentation difficulty of this regional knowledge image is smaller, and then the corresponding character discrimination coefficient is also larger. Through the specific analysis of the character discrimination coefficient, it helps to more accurately distinguish each initially recognized character, so that the segmentation of each initially recognized character is more accurate, and further improves the accuracy of the character recognition and extraction of the handwritten documents in the knowledge base of power enterprises.
[0043] Specifically, the character segmentation weight factor is obtained from a preset database, and the character segmentation weight factor represents the influence degree of the character segmentation data on the character discrimination coefficient. Each character segmentation data has a unique mapping relationship with the character discrimination coefficient, and the value range is from 0 to 1; in a specific embodiment, the character horizontal weight factor, character vertical weight factor, and character edge weight factor are obtained in real time from a preset database, which respectively represent the influence degrees of the character horizontal spacing, character vertical spacing, and character edge density on the character discrimination coefficient, and the sum of the three is 1.
[0044] Specifically, the character distinction threshold is obtained from a preset database. In a specific embodiment, the character segmentation data related to handwritten documents in the historical data is substituted into the specific restriction expression of the character distinction coefficient to obtain the corresponding data set, and the character distinction threshold is obtained by performing a mean operation on the data set.
[0045] Specifically, the specific expression of the sigmoid function is as follows:
[0046]
[0047] Here, x represents the input of the function and y represents the output of the function.
[0048] Furthermore, the specific steps of judging whether to perform manual verification by the ratio of single-character pieces are as follows: A1, comparing the ratio of single-character pieces with the preset ratio of single-character pieces. If the ratio of single-character pieces is less than the preset ratio, manual verification is performed, otherwise A2 is executed. Manual verification means feeding back the corresponding character piece to the preset verification personnel to obtain the characters of the corresponding character piece. The ratio of single-character pieces indicates the ratio of the number of single-character pieces to the total number of character pieces. A2, performing character morphology evaluation on each character piece to obtain character morphology data. If the character piece does not have character morphology data, the corresponding character piece is deleted. Otherwise, A3 is executed. Character morphology evaluation is used to judge the integrity of the character according to the character morphology. A3, if the character morphology data of each character piece is not less than the corresponding reference morphology data, manual verification is not performed. Otherwise, the character piece is combined with the horizontally adjacent character piece to obtain a non-single character piece. The reference morphology data includes the maximum character confidence value and the maximum character connectivity value.
[0049] In this embodiment, the preset single-character piece ratio is pre-set by the preset staff according to the specific segmentation of the regional knowledge picture, and is stored in the preset database, and is obtained in real time from the preset database when used; the method provided by this embodiment not only ensures the integrity of the characters in each character piece, but also removes the altered characters caused by typos, thereby reducing interference for subsequent character comparison and recognition, and when it is detected that the ratio of single-character pieces is lower than the preset single-character piece ratio, it indicates that the corresponding handwritten document may have a high degree of connected strokes, and manual verification is directly performed, thereby improving the efficiency of character comparison, and further improving the accuracy of character extraction of handwritten documents in the electric power enterprise knowledge base.
[0050] Specifically, the reference morphological data is obtained from a preset database. In a specific embodiment, the reference morphological data is set by a preset staff according to the specific conditions of the regional knowledge image, and is obtained from the preset database in real time when used.
[0051] Further, character recognition analysis is performed based on the character recognition data to obtain a set of words to be compared. The specific process is as follows: The character slices are numbered and compared with a general character library, and a character set is obtained. The general character library represents a general database applicable to all databases for character recognition; irrelevant part-of-speech recognition is performed on the character set, and the irrelevant words in the character slices are removed to obtain a set of characters to be compared and character recognition data. The character recognition coefficient is obtained based on the character recognition data and the word correction factor; if the character recognition coefficient is not less than the character recognition threshold in the preset database, the character slices in the set of characters to be compared are combined according to the character slice numbers to obtain the corresponding set of words to be compared, otherwise the character discrimination threshold is optimized; if the character recognition coefficient corresponding to the character slice obtained based on the optimized character discrimination threshold is less than the character recognition threshold, manual verification is performed, otherwise the corresponding set of words to be compared is obtained.
[0052] In this embodiment, the irrelevant part-of-speech recognition is implemented through an NLP large model; optimizing the character discrimination threshold means reducing the character discrimination threshold by a preset ratio. If the character recognition coefficient is still less than the character recognition threshold after the character discrimination threshold is reduced to the lowest character area threshold, manual verification is directly performed; through the method provided in this embodiment, it helps to more comprehensively analyze the character comparison situation, and then make corresponding adjustments in a timely manner, thereby simply improving the accuracy of character recognition for the updated handwritten documents in the power enterprise knowledge base.
[0053] Specifically, the character recognition threshold is obtained from the preset database. In a specific embodiment, the character recognition data related to the handwritten documents in the historical data is substituted into the specific limit expression of the character recognition coefficient to obtain a corresponding data set, and the mean value operation is performed on the data set to obtain the character recognition threshold.
[0054] Further, the specific process of obtaining the character recognition coefficient is as follows: Obtain the character recognition weight factor and the maximum matching time from the preset database. The character recognition weight factor includes a character matching weight factor, a character overlap weight factor, a character matching time weight factor, and a mis-matching rate weight factor; the word correction difference (i.e., 1 - CCF) is obtained according to the word correction factor, and the word correction difference represents the difference between the word correction factor and the value 1; after performing character recognition processing on each character recognition data respectively and then performing weighted processing in combination with the character matching weight factor, a character recognition processing result is obtained. The character recognition processing is used to unify the logical relationship between the character recognition data and the character recognition coefficient; the character recognition processing result is multiplied by the word correction difference to obtain the character recognition coefficient. The character recognition coefficient represents the quantitative data of the accuracy degree of character recognition after character segmentation jointly by the character matching rate, the character overlap rate, the character matching completion time, and the mis-matching rate;
[0055] The specific limit expression of the character recognition coefficient is as follows:
[0056]
[0057] In the formula, CMR represents the character matching rate, COR represents the character overlap rate, CMT represents the character matching completion time, CMT0 represents the maximum matching time, MR represents the mismatch rate, and γ a Represents the character matching weight factor, γ b represents the character overlap weight factor, γ c Represents the character matching time weight factor, γ d represents the mismatch rate weight factor, CCF represents the word correction factor, and CRC represents the character recognition coefficient.
[0058] In this embodiment, the character recognition coefficient algorithm combines character recognition data, character recognition weight factor, maximum matching time and word correction factor for comprehensive analysis to obtain the character recognition coefficient, wherein, as the character matching rate increases, it means that there are more character segments that are successfully matched, and the corresponding character recognition coefficient also increases accordingly. As the character overlap rate, character matching completion time and mismatch rate increase, it means that the probability of character matching errors is greater, and the corresponding character recognition coefficient decreases instead. When the word correction factor is smaller, it means that the degree of character matching that needs to be corrected due to morphological analysis and word ambiguity is lower, and the corresponding character recognition coefficient is larger. Through detailed analysis of the character recognition coefficient, the situation of character matching is understood in more detail, and the character recognition data is obtained based on the comparison, and the overall comparison situation is analyzed, so that timely optimization is performed to obtain a more accurate set of words to be matched, which helps the NLP large model to understand the semantics of updated handwritten documents and improves the accuracy of the NLP large model in entity recognition in updated handwritten documents.
[0059] Specifically, the maximum matching time is obtained from a preset database, indicating the time set by the system to limit the comparison of the character set to be compared with the universal character library. In a specific embodiment, the maximum matching time is set by a preset staff according to the actual situation of the specific uploaded updated handwritten document.
[0060] Specifically, the character recognition weight factor is obtained from a preset database, and the character recognition weight factor represents the degree of influence of the character recognition data on the character recognition coefficient. Each character recognition data is uniquely mapped to the character recognition coefficient, and the value range is 0 to 1. In a specific embodiment, the character matching weight factor, the character overlap weight factor, the character matching time weight factor, and the mismatch rate weight factor are obtained from the preset database in real time, respectively representing the degree of influence of the character matching rate, the character overlap rate, the character matching completion time, and the mismatch rate on the character recognition coefficient, and the sum of the four is 1.
[0061] Further, the specific process of obtaining the word correction factor is as follows: Obtain the correction data and the correction weights in the preset database. The correction data includes morphological similarity and ambiguity, and the correction weights include morphological weight and ambiguity weight. Input the correction data and the correction weights into the correction model to obtain the corresponding word correction factor. The word correction factor represents the quantitative data of the combined effect of morphological similarity and ambiguity on reducing the matching compliance degree of semantic comparison pairs. The correction model is used to construct the logical relationship between the correction data and the word correction factor.
[0062] The specific constraint expression of the correction model is:
[0063]
[0064] In the formula, MS represents morphological similarity, AD represents ambiguity, α represents morphological weight, β represents ambiguity weight, and CCF represents the word correction factor.
[0065] In this embodiment, the hyperbolic tangent function is used to normalize the word correction factor. When the morphological similarity and ambiguity are higher, it indicates that the degree of correction required for the character recognition coefficient and the matching coefficient of the power enterprise word library is higher, and the corresponding word correction factor is also larger. The word correction factor obtained through the correction model is applied to the specific analysis of the character recognition coefficient and the matching coefficient of the power enterprise word library, correcting the influence of character morphological analysis and word semantic ambiguity on the matching compliance degree of word comparison, thereby making the analysis of the character recognition coefficient and the matching coefficient of the power enterprise word library more accurate.
[0066] Specifically, morphological similarity is used to measure the similarity between two morphologies (such as characters, graphics, or regions), and is obtained through morphological analysis. Ambiguity represents the possibility of ambiguity in a word during the word matching process, and is calculated through information entropy. Information entropy is an index to measure information uncertainty, and the information entropy is calculated through the information entropy formula. The specific information entropy formula is as follows:
[0067]
[0068] It should be explained that H(X) represents information entropy, P(x i ) represents the occurrence probability of an event, i represents the event number, and D represents the total number of events. Here, an event represents different meanings of a word. For example, if a word has multiple meanings and the occurrence probability of each meaning is close, then the ambiguity of the word is higher.
[0069] Specifically, the correction weight is obtained from a preset database, and the correction weight represents the influence degree of correction data on the word correction factor. There is a unique mapping relationship between each correction data and the word correction factor, and the value range is from 0 to 1; in a specific embodiment, the morphological weight and the ambiguity weight are obtained in real time from the preset database, which respectively represent the influence degrees of morphological similarity and ambiguity on the word correction factor, and the sum of the two is 1.
[0070] Further, the word set to be compared is compared with the knowledge word library of the power enterprise, and the matching degree analysis is carried out according to the word library data obtained after the comparison. The specific process is as follows: obtain the word library matching weight factor and the maximum completion time from the preset database. The word library matching weight factor includes the context matching degree weight factor, the word library matching weight factor, the word library completion time weight factor, and the word library word null value weight factor; after performing word library comparison processing on each word library data respectively, weighted processing is carried out in combination with the word library matching weight factor to obtain the word library comparison processing result. The word library comparison processing is used to establish the logical relationship between the word library data and the matching coefficient of the power enterprise word library; the word library comparison processing result is operated with the word correction difference to obtain the power enterprise word library matching coefficient. The power enterprise word library matching coefficient represents the quantitative data of the degree of conformity of the recognized characters with the vocabulary in the professional knowledge library of the power enterprise jointly by the context matching degree, the word library matching success rate, the word library matching completion time, and the word library word null value rate.
[0071] The specific limiting expression of the power enterprise word library matching coefficient is as follows:
[0072]
[0073] In the formula, CMD represents the context matching degree, SLM represents the word library matching success rate, LMT represents the word library matching completion time, LMT0 represents the maximum completion time, EV represents the word library word null value rate, δ a represents the context matching degree weight factor, δ b represents the word library matching weight factor, δ c represents the word library completion time weight factor, δ d represents the word library word null value weight factor, CCF represents the word correction factor, and PEV represents the power enterprise word library matching coefficient.
[0074] In this embodiment, the power enterprise thesaurus matching coefficient algorithm comprehensively analyzes the thesaurus data, the thesaurus matching weight factor, the maximum completion time, and the word correction factor to obtain the power enterprise thesaurus matching coefficient. In the formula, as the context matching degree and the thesaurus matching success rate increase, more character slices indicating the successful comparison of the word set to be compared are obtained, and the corresponding power enterprise thesaurus matching coefficient also increases. As the thesaurus matching completion time and the thesaurus word null value rate increase, the probability of errors in the comparison of the word set to be compared increases, and the corresponding power enterprise thesaurus matching coefficient decreases instead. When the word correction factor is smaller, it indicates that the character comparison degree of the word set to be compared that needs to be corrected due to morphological analysis and word ambiguity is lower, and the corresponding power enterprise thesaurus matching coefficient is larger. Through the detailed analysis of the power enterprise thesaurus matching coefficient, the comparison situation of the word set to be compared is understood in more detail. Moreover, the acquisition of the power enterprise thesaurus matching coefficient is based on the overall analysis after the comparison, so as to optimize it in a timely manner to obtain a more accurate text report, improve the semantic understanding degree of the NLP large model for the updated handwritten document, and at the same time improve the accuracy of entity recognition in the updated handwritten document by the NLP large model.
[0075] It should be understood that the character recognition coefficient and the power enterprise thesaurus matching coefficient are obtained and analyzed without numbering the character slices, which is different from the character discrimination coefficient that numbers the initially recognized character slices. The reason is that the character recognition coefficient and the power enterprise thesaurus matching coefficient are based on the overall analysis of the comparison situation after the comparison, so numbering is not required. The character discrimination coefficient is used to split the initially recognized character slices and needs to analyze the initially recognized character slices separately, so numbers are introduced for distinction.
[0076] Specifically, the maximum completion time is obtained from the preset database and represents the time limit set by the system for the comparison of the word set to be compared with the power enterprise knowledge thesaurus. In a specific embodiment, the maximum completion time is set by the preset staff according to the actual situation of the specific uploaded updated handwritten document.
[0077] Specifically, the thesaurus matching weight factor is obtained from the preset database, and the thesaurus matching weight factor represents the influence degree of the thesaurus data on the power enterprise thesaurus matching coefficient. There is a unique mapping relationship between each thesaurus data and the power enterprise thesaurus matching coefficient, and the value range is from 0 to 1. In a specific embodiment, the context matching degree weight factor, the thesaurus matching weight factor, the thesaurus completion time weight factor, and the thesaurus word null value weight factor are obtained in real time from the preset database, which respectively represent the influence degrees of the context matching degree, the thesaurus matching success rate, the thesaurus matching completion time, and the thesaurus word null value rate on the power enterprise thesaurus matching coefficient, and the sum of the four is 1.
[0078] Further, the specific process of determining whether to conduct manual review is as follows: Compare the matching coefficient of the power enterprise thesaurus with the power thesaurus threshold obtained from the preset database. If the matching coefficient of the power enterprise thesaurus is not less than the power thesaurus threshold, then combine all character slices in the order of their numbers according to the comparison result of the thesaurus to be compared, and at the same time automatically generate a text report based on the NLP large model for entity recognition. If the matching coefficient of the power enterprise thesaurus is less than the power thesaurus threshold, then feedback the character slices in the thesaurus to be compared without comparison results to the preset staff for manual verification.
[0079] In this embodiment, by comparing with the power enterprise knowledge thesaurus, it ensures the professionalism of the recognized handwritten document-related characters in the field of power enterprise knowledge, and is also beneficial to the accuracy of semantic understanding when the NLP large model automatically generates a text report and conducts entity recognition, thereby improving the efficiency and accuracy of updating and maintaining the power enterprise knowledge base based on the NLP large model. And introducing manual verification for assistance further ensures the accuracy of extracting handwritten document information.
[0080] Specifically, the power thesaurus threshold is obtained from the preset database. In a specific embodiment, substitute the thesaurus data related to handwritten documents in historical data into the specific limit expression of the matching coefficient of the power enterprise thesaurus to obtain the corresponding data set, and perform a mean operation on the data set to obtain the power thesaurus threshold.
[0081] The method for updating and maintaining the power enterprise knowledge base based on the NLP large model provided in the embodiment of the present application includes the following steps: Perform image processing on the regional knowledge image to obtain the corresponding character segmentation data, and perform segmentation processing on the regional knowledge image based on the character segmentation data to obtain character slices, where the character slices include single-character slices and non-single-character slices. Determine whether to conduct manual verification based on the proportion of single-character slices. If no manual verification is conducted, then evaluate the character morphology of each character slice to obtain character morphology data, and accordingly judge the integrity of each character slice. Compare the character slices with the general character library to obtain the character set to be compared and character recognition data, and perform character recognition analysis based on the character recognition data to obtain the word set to be compared. Compare the word set to be compared with the power enterprise knowledge thesaurus, and perform a matching degree analysis based on the thesaurus data obtained after comparison to determine whether to conduct manual review.
[0082] In this embodiment, the method provided in the embodiment of the present application not only improves the accuracy of extracting handwritten information that needs to be updated in the power enterprise knowledge base, but also is beneficial to the efficiency and accuracy of updating and maintaining the power enterprise knowledge base by the NLP large model.
[0083] In summary, in the embodiment of the present application, character slices are obtained by segmenting regional knowledge pictures. Then, it is judged whether manual verification is required based on the proportion of single character slices. Next, the character slices are compared with the general character library to obtain a set of characters to be compared, and character recognition analysis is performed to obtain a set of words to be compared. Finally, the set of words to be compared is compared with the power enterprise knowledge word library to perform matching degree analysis and judge whether manual review is required, so as to obtain more accurate extracted handwritten information, and further improve the accuracy of extracting handwritten information that needs to be updated in the power enterprise knowledge base, effectively solving the problem of inaccurate extraction of handwritten information that needs to be updated in the power enterprise knowledge base in the prior art.
[0084] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and the instructions executed on the computer or other programmable device provide for implementing the specified functions in Figure 1steps of one process or multiple processes and / or boxes Figure 1 steps of functions specified in one box or multiple boxes.
[0088] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0089] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A power enterprise knowledge base update and maintenance system based on the NLP large model, characterized in that, It includes a knowledge picture segmentation module, a character integrity judgment module, a general library comparison module, and a power vocabulary library comparison module: Among them, the knowledge picture segmentation module is used to perform picture processing on the regional knowledge picture to obtain corresponding character segmentation data, and perform segmentation processing on the regional knowledge picture based on the character segmentation data to obtain character slices, and the character slices include single-character slices and non-single-character slices; The character integrity judgment module is used to judge whether to perform manual verification through the single-character slice ratio. If no manual verification is performed, it evaluates the character form of each character slice to obtain character form data, and accordingly judges the integrity of each character slice; The general library comparison module is used to compare the character slices with the general character library to obtain a set of characters to be compared and character recognition data, and perform character recognition analysis based on the character recognition data to obtain a set of words to be compared; The power vocabulary library comparison module is used to compare the set of words to be compared with the power enterprise knowledge vocabulary library, and perform a matching degree analysis based on the vocabulary data obtained after the comparison to judge whether to perform manual review.
2. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 1, characterized in that: The character segmentation data includes character horizontal spacing, character vertical spacing, and character edge density; The character form data includes character confidence and character connectivity; The character recognition data includes character matching rate, character overlap rate, character matching completion time, and mis-matching rate; The character overlap rate represents the success rate of evaluating the match by analyzing the overlapping part of the shape of the character and the character in the library; The vocabulary data includes context matching degree, vocabulary library matching success rate, vocabulary library matching completion time, and vocabulary library word null value rate; The context matching degree represents the matching degree of the character with the context of its preceding and following characters; The vocabulary library word null value rate represents the proportion of words that are not matched in the vocabulary library in the total vocabulary.
3. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 2, characterized in that: The specific process of performing segmentation processing on the regional knowledge picture based on the character segmentation data to obtain character slices is as follows: Perform data normalization processing on the character segmentation data obtained by performing picture processing on the regional knowledge picture, and obtain the character segmentation weight factor and the minimum reference ratio from the preset database. The character segmentation weight factor includes a character horizontal weight factor, a character vertical weight factor, and a character edge weight factor; Obtain a segmentation difficulty quantization coefficient based on the character horizontal spacing, the character vertical spacing, and the minimum reference ratio. The segmentation difficulty quantization coefficient represents the quantization data of the combined influence of the character horizontal spacing and the character vertical spacing on the difficulty of segmenting the handwritten document; Perform preliminary recognition on the regional knowledge picture to obtain preliminary recognition characters and number them, and obtain the character segmentation data of each preliminary recognition character; Perform weighted processing on the character segmentation data in combination with the character segmentation weight factor and then perform character differentiation processing to obtain a character differentiation processing result. Perform arithmetic processing on the character differentiation processing result and the segmentation difficulty quantization coefficient to obtain a character differentiation coefficient. The character differentiation processing is used to establish a data mapping relationship between the character segmentation data and the character differentiation coefficient. The character differentiation coefficient represents the quantization data of the combined influence of the character horizontal spacing, the character vertical spacing, and the character edge density on the independence between characters; If the character discrimination coefficient is not less than the character discrimination threshold in the preset database, the regional knowledge picture is segmented into corresponding column character slices by abnormal column segmentation; otherwise, the character segmentation data of the next horizontally adjacent initially recognized character is obtained until the corresponding character discrimination coefficient is not less than the character discrimination threshold for abnormal column segmentation. Based on the character discrimination coefficients of the initially recognized characters in each column character slice, row segmentation is performed to obtain character slices.
4. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 1, characterized in that: The specific steps for determining whether to perform manual verification by the single character slice ratio are as follows: A1. Compare the single character slice ratio with the preset single character slice ratio. If the single character slice ratio is less than the preset single character slice ratio, manual verification is performed; otherwise, A2 is executed. Manual verification means feeding the corresponding character slice back to the preset verification personnel to obtain the characters of the corresponding character slice. A2. Perform character form evaluation on each character slice to obtain character form data. If the character slice does not have character form data, the corresponding character slice is deleted; otherwise, A3 is executed. The character form evaluation is used to judge the integrity of the character according to the form of the character. A3. If the character form data of each character slice is not less than the corresponding reference form data, no manual verification is performed; otherwise, the character slice is combined with the horizontally adjacent character slice to obtain a non-single character slice. The reference form data includes the maximum character confidence and the maximum character connectivity.
5. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 1, characterized in that: The specific process of performing character recognition analysis based on the character recognition data to obtain the word set to be compared is as follows: Number the character slices and compare them with the general character library to obtain a character set. Perform irrelevant part-of-speech recognition on the character set, remove the irrelevant words in the character slices to obtain the word set to be compared and character recognition data, and obtain the character recognition coefficient according to the character recognition data and the word correction factor. If the character recognition coefficient is not less than the character recognition threshold in the preset database, the character slices in the word set to be compared are combined according to the character slice numbers to obtain the corresponding word set to be compared; otherwise, the character discrimination threshold is optimized. If the character recognition coefficient corresponding to the character slice obtained based on the optimized character discrimination threshold is less than the character recognition threshold, manual verification is performed; otherwise, the corresponding word set to be compared is obtained.
6. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 5, characterized in that: The specific process of obtaining the character recognition coefficient is as follows: Obtain the character recognition weight factor and the maximum matching time from the preset database. The character recognition weight factor includes the character matching weight factor, the character overlap weight factor, the character matching time weight factor, and the mis-matching rate weight factor. Obtain the word correction difference according to the word correction factor. After performing character recognition processing on each character recognition data respectively and then performing weighted processing in combination with the character matching weight factor to obtain the character recognition processing result. The character recognition processing is used to unify the logical relationship between the character recognition data and the character recognition coefficient. Perform a product operation on the character recognition processing result and the word correction difference to obtain the character recognition coefficient. The character recognition coefficient is used to represent the quantitative data of the accuracy degree of character recognition after character segmentation jointly by the character matching rate, the character overlap rate, the character matching completion time, and the mis-matching rate.
7. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 6, characterized in that: The specific process of obtaining the word correction factor is as follows: Obtaining correction data and correction weights in a preset database, wherein the correction data includes morphological similarity and ambiguity, and the correction weights include morphological weights and ambiguity weights; The correction data and correction weight are input into the correction model to obtain the corresponding word correction factor. The word correction factor represents the quantitative data of the degree of matching of the semantic comparison reduced by the morphological similarity and ambiguity. The correction model is used to construct a logical relationship between the correction data and the word correction factor.
8. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 2, characterized in that: The word set to be compared is compared with the electric power enterprise knowledge word library, and the matching degree analysis is performed based on the word library data obtained after the comparison. The specific process is as follows: Obtaining a vocabulary matching weight factor and a maximum completion time from a preset database, wherein the vocabulary matching weight factor includes a context matching weight factor, a vocabulary matching weight factor, a vocabulary completion time weight factor, and a vocabulary word null value weight factor; After performing a vocabulary comparison process on each vocabulary data, a weighted process is performed in combination with a vocabulary matching weight factor to obtain a vocabulary comparison process result. The vocabulary comparison process is used to establish a logical relationship between the vocabulary data and the power enterprise vocabulary matching coefficient; The result of the vocabulary comparison processing and the word correction difference are processed to obtain the electric power enterprise vocabulary matching coefficient, which is used to express the quantitative data of the degree of conformity between the recognized characters and the vocabulary in the electric power enterprise professional knowledge base vocabulary, including the context matching degree, vocabulary matching success rate, vocabulary matching completion time and vocabulary word empty value rate.
9. The power enterprise knowledge base update and maintenance system based on the NLP large model according to claim 8, characterized in that: The specific process of determining whether to conduct manual review is as follows: Compare the power enterprise vocabulary matching coefficient with the power vocabulary threshold obtained from the preset database: If the power enterprise vocabulary matching coefficient is not less than the power vocabulary threshold, all character segments are combined in numerical order according to the comparison results of the vocabulary to be compared, and a text report is automatically generated based on the NLP large model for entity recognition; If the power enterprise vocabulary matching coefficient is less than the power vocabulary threshold, the character segments without matching results in the vocabulary to be compared will be fed back to the preset staff for manual verification.
10. A method for updating and maintaining a knowledge base of power enterprises based on large NLP models, characterized in that: The following steps are involved: Performing image processing on the regional knowledge image to obtain corresponding character segmentation data, and segmenting the regional knowledge image based on the character segmentation data to obtain character slices, wherein the character slices include single-character slices and non-single-character slices; The ratio of a single character slice is used to determine whether manual verification is required. If manual verification is not required, the character morphology of each character slice is evaluated to obtain character morphology data, and the integrity of each character slice is determined based on this data. Compare the character slice with the universal character library to obtain a character set to be compared and character recognition data, perform character recognition analysis based on the character recognition data, and obtain a word set to be compared; The word set to be compared is compared with the knowledge word library of the electric power enterprise, and the matching degree analysis is performed based on the word library data obtained after the comparison to determine whether manual review is required.
Citation Information
Patent Citations
Natural language processing method, device, equipment and storage medium
CN114428788B
A knowledge enhancement method and system for large language models
CN118779470B