Military document information extraction and knowledge graph construction method and system
By combining text layout features, term matching algorithms and military term knowledge bases, we can identify real text areas in military documents; use symbol shape feature analysis and military symbol knowledge bases to identify symbol annotations; and build a knowledge graph through mapping relationships, the problem that traditional technology is difficult to identify military terms and symbols is solved, and efficient and accurate information extraction and structured processing are achieved.
Patent Information
- Application Number
- CN202510302128.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
AI Technical Summary
In the PDF processing of military documents, traditional OCR technology is difficult to accurately identify and extract military terms and symbol annotations, especially in complex text layouts and high-density text environments.
By combining text layout features, term matching algorithms and military term knowledge bases, we can accurately identify real text areas; using symbol shape feature analysis, pattern recognition algorithms and military symbol knowledge bases, we can accurately identify symbol annotations; and by constructing the mapping relationship between text and symbol annotations, we can correct the recognition results and generate structured military element description data.
It improves the accuracy and efficiency of military document information extraction, ensures the accurate correlation between text content and symbol annotations, maintains the integrity and consistency of information, and builds a military element knowledge graph to support information retrieval, analysis and decision-making in the military field.
Smart Images

Figure CN120124732A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of information extraction, and particularly to a method and system for military document information extraction and knowledge graph construction. Background Art
[0002] In the processing of PDF files in the military field, the extraction of key texts faces a unique technical problem: how to accurately identify and extract military terms and symbol annotations in a specific format from complex military documents. Military documents usually contain a large amount of mixed Chinese and English texts, and the text layout is complex, with texts and symbol annotations intertwined. This complexity causes traditional OCR technologies to be prone to recognition errors when processing these documents, especially when dealing with military terms and symbols.
[0003] First of all, when a PDF document is converted into an image, due to the special layout of military documents, the image quality may be affected. Especially when dealing with high-density texts and complex symbols, the texts and symbols in the image may be blurred or distorted. This distortion will directly affect the subsequent detection of text and symbol areas, resulting in inaccurate detection areas and thus affecting the recognition effect.
[0004] Secondly, in the text and symbol area detection process, the symbol annotations in military documents are often closely adjacent to or even overlap with the texts. Traditional area detection algorithms are difficult to accurately distinguish these areas, and are prone to missed detection or false detection. Especially when dealing with military symbols, the shapes and positions of the symbols are variable, increasing the difficulty of detection.
[0005] In the text and symbol content recognition process, the recognition of military terms is particularly crucial. Military terms often have specific abbreviations, symbols, and formats, and traditional OCR models are prone to misrecognition when dealing with these terms. In addition, the symbol annotations in military documents often contain specific military meanings, and recognizing these symbols requires combining with a specific military knowledge base, while existing OCR models lack this deep combination ability.
[0006] Finally, in the information structuring and storage process, the extracted text and symbol content need to be structured for further analysis. However, the information in military documents is often scattered in multiple areas, and the relationships between the information are complex. How to effectively structure this information and extract key information is a technical difficulty. Especially when dealing with military terms and symbols, how to accurately associate this information to ensure the integrity and consistency of the information is an urgent problem to be solved. Summary of the Invention
[0007] An embodiment of the present invention provides a method and system for military document information extraction and knowledge graph construction, which can accurately identify and associate text with symbolic annotations, correct the recognition results, generate structured military element description data, and construct a military element knowledge graph, thereby improving the accuracy and efficiency of military document information processing.
[0008] To achieve the above object, in a first aspect, the present invention provides a method and system for military document information extraction and knowledge graph construction, wherein the method includes: obtaining a military document PDF file, dividing the text area according to the text layout features, and generating a set of candidate text areas. For the set of candidate text areas, a term matching algorithm is used to match with the military term knowledge base to screen out the real text areas. For non-text areas, shape analysis is performed according to the symbol shape features to generate a set of candidate symbol annotation areas. For the set of candidate symbol annotation areas, a pattern recognition algorithm is used to match with the military symbol knowledge base to screen out the real symbol annotation areas. According to the relative positions of the real text areas and the real symbol annotation areas, the relevance between the two is determined, and a mapping relationship between the text and the symbol annotations is constructed. For the text areas with the mapping relationship established, an OCR model is used to recognize the text content, and the recognition result is corrected according to the military term knowledge base. For the symbol annotation areas with the mapping relationship established, a symbol recognition model is used to obtain the symbol annotation information, and the recognition result is corrected according to the military symbol knowledge base. According to the mapping relationship between the text and the symbol annotations, the corrected text recognition result and the symbol annotation information are associated to generate structured military element description data, and a military element knowledge graph is constructed.
[0009] Second aspect, the present invention provides a military document information extraction and knowledge graph construction system, including: a candidate text region generation module, a real text region screening module, a candidate symbol annotation region generation module, a real symbol annotation region screening module, a mapping relationship construction module, a text region correction module, a symbol annotation region correction module, and a knowledge graph construction module. The candidate text region generation module is used to obtain a military document PDF file, divide text regions according to text layout features, and generate a set of candidate text regions. The real text region screening module is used to match the set of candidate text regions with a military term knowledge base by using a term matching algorithm, and screen out real text regions. The candidate symbol annotation region generation module is used to perform shape analysis on non-text regions according to symbol shape features, and generate a set of candidate symbol annotation regions. The real symbol annotation region screening module is used to match the set of candidate symbol annotation regions with a military symbol knowledge base by using a pattern recognition algorithm, and screen out real symbol annotation regions. The mapping relationship construction module is used to determine the relevance between the real text region and the real symbol annotation region according to their relative positions, and construct a mapping relationship between text and symbol annotation. The text region correction module is used to recognize the text content of the text region with a mapping relationship established by using an OCR model, and correct the recognition result according to the military term knowledge base. The symbol annotation region correction module is used to obtain symbol annotation information for the symbol annotation region with a mapping relationship established by using a symbol recognition model, and correct the recognition result according to the military symbol knowledge base. The knowledge graph construction module is used to associate the corrected text recognition result and symbol annotation information according to the mapping relationship between text and symbol annotation, generate structured military element description data, and construct a military element knowledge graph.
[0010] Third aspect, the present invention provides an electronic device, including: at least one processor; and
[0011] a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned military document information extraction and knowledge graph construction method.
[0012] Fourth aspect, the present invention provides a computer-readable storage medium, including a computer program and instructions, when the computer program or instructions run on a computer, enabling the computer to execute the above-mentioned military document information extraction and knowledge graph construction method.
[0013] Compared with the prior art, the military document information extraction and knowledge graph construction method and system according to the present invention have the following beneficial effects:
[0014] 1. By combining text layout features, term matching algorithms, and a military term knowledge base, the present invention can more accurately identify real text regions in military documents, reducing the misrecognition rate of traditional OCR technology when processing complex military documents; by analyzing symbol shape features, pattern recognition algorithms, and a military symbol knowledge base, the present invention can accurately identify and extract symbol annotations in military documents, improving the accuracy of symbol recognition.
[0015] 2. By constructing a mapping relationship between text and symbol annotations, the present invention ensures an accurate association between text content and symbol annotations, which helps maintain the integrity and consistency of information in subsequent information structuring processes.
[0016] 3. The present invention uses structuring processing technology to effectively structure the extracted text and symbol content, generating military element description data that is easy to further analyze; through knowledge inference algorithms, the present invention can extract relevant information from the military knowledge base for comparison, further enriching the description data of military elements and inferring the effectiveness attributes of military elements.
[0017] 4. The present invention finally constructs a military element knowledge graph, providing strong support for information retrieval, analysis, and decision-making in the military field.
[0018] 5. The present invention realizes the automated process of military document information extraction and knowledge graph construction, greatly improving work efficiency and reducing manual intervention and errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flowchart of a method for extracting military document information and constructing a knowledge graph in Embodiment 1 of the present invention;
[0020] Figure 2 is a schematic structural diagram of a system for extracting military document information and constructing a knowledge graph in Embodiment 2 of the present invention;
[0021] Figure 3 is a schematic structural diagram of an electronic device in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The following further elaborates on the embodiments of the present invention in conjunction with the drawings and examples. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention, rather than limiting the embodiments of the present invention. Additionally, it should be noted that for ease of description, only parts related to the embodiments of the present invention are shown in the drawings, rather than all structures.
[0023] For ease of understanding, the main implementation concepts of each embodiment of the present invention are briefly described first.
[0024] In the field of military document processing, as a common document format, PDF files face many technical challenges in information extraction and knowledge graph construction.
[0025] Military documents usually contain a large amount of mixed Chinese and English text, and the text layout is complex. This complexity makes it difficult to directly apply traditional text processing techniques. When a PDF document is converted into an image for processing, due to the special layout of military documents (such as high-density text, complex symbols, etc.), the image quality may be affected, resulting in blurring or distortion. This distortion will directly affect the subsequent detection of text and symbol regions, leading to inaccurate detection regions and thus affecting the recognition effect.
[0026] Symbol annotations in military documents are often closely adjacent to or even overlap with the text. Traditional region detection algorithms are difficult to accurately distinguish these regions, and it is easy to miss detections or make false detections. The shapes and positions of military symbols are variable, increasing the difficulty of detection. Traditional algorithms may not be able to adapt to this variability, resulting in a decrease in recognition accuracy.
[0027] Military terms often have specific abbreviations, symbols, and formats. When traditional OCR (Optical Character Recognition) models process these terms, they are prone to misrecognition. This is because traditional OCR models may not be trained and optimized for the particularity of military terms. Symbol annotations in military documents often contain specific military meanings, and identifying these symbols requires combining a specific military knowledge base. However, existing OCR models lack this deep combination ability, resulting in low recognition accuracy for symbol annotations.
[0028] The information in military documents is often scattered in multiple regions, and the relationships between the information are complex. How to effectively structure this information and extract key information is a technical difficulty. When dealing with military terms and symbols, how to accurately associate this information to ensure the integrity and consistency of the information is an urgent problem to be solved. Traditional information structuring methods may not be able to handle this complex relationship and correlation.
[0029] In summary, military document information extraction and knowledge graph construction face many technical challenges. To overcome these challenges, a method and system capable of accurately identifying and extracting military terms and symbol annotations in a specific format need to be developed.
[0030] By discovering the defects in the background technology as described above, the inventors provide a method and system for military document information extraction and knowledge graph construction. By combining various technical means such as text layout features, term matching algorithms, pattern recognition algorithms, and military knowledge bases, the present invention can effectively improve the accuracy and efficiency of military document information extraction, providing strong support for the informatization and intelligence of the military field.
[0031] Example 1Figure 1 It is a schematic flowchart of a method for extracting military document information and constructing a knowledge graph in the first embodiment of the present invention. As Figure 1 shown, the first embodiment provides a method for extracting military document information and constructing a knowledge graph, including:
[0032] Step S100: Obtain a military document PDF file, divide the text area according to text layout features, and generate a set of candidate text areas;
[0033] Specifically, the system first obtains PDF documents in the military field. Such documents often contain a variety of elements such as complex Chinese-English mixed texts and symbol annotations, and have diverse layouts. To effectively process these complex documents, the system divides the area according to the layout features of the text. This includes analyzing various information such as the line spacing, alignment, font size, and color of the text. For example, if the line spacing of the text is less than a preset threshold, the system determines that these texts belong to the same text area; if the difference in font size exceeds the preset range, these texts are divided into different text areas. In addition, the system will also classify the text content of the same color into the same candidate area according to the font color feature. After completing the preliminary area division, the system will adopt a rule-based area merging algorithm to merge adjacent candidate text areas with the same alignment to further optimize the division of the text area. Finally, through the preset layout feature combination rules, the system will eliminate the areas that do not meet the requirements, thereby generating a more accurate and complete set of candidate text areas.
[0034] Step S200: For the set of candidate text areas, use a term matching algorithm to match with the military term knowledge base, and screen out the real text areas;
[0035] Specifically, the military term knowledge base contains a wealth of information such as a large number of professional terms, standard names, synonyms, hyponymy relations, and English translations in the military field. The algorithm first extracts the text content in the candidate text area to form a text feature set, and then constructs a term matching rule based on the military term knowledge base. Then, the algorithm compares the text feature set with the term matching rule item by item and calculates the matching degree of each item. If the text content of a certain candidate text area has a matching degree higher than the preset threshold with a certain term in the military term knowledge base, then this area is determined to be a real text area. This process effectively screens out the text areas containing important military information, excludes possible misidentifications or irrelevant texts, and thus improves the accuracy and efficiency of military document information extraction. Through this method, even in the face of the complex and changeable text layout and a large number of professional terms in military documents, the system can accurately identify key information, laying a solid foundation for subsequent information structuring processing and knowledge graph construction.
[0036] Step S300: For non-text regions, perform shape analysis based on symbol shape features to generate a set of candidate symbol annotation regions.
[0037] Specifically, first obtain the image data of non-text regions, which may contain various complex military symbols or other non-verbal information. Then, use feature extraction algorithms to deeply analyze this image data, mainly focusing on key features such as the shape complexity of the symbols, line thickness, and filling color. These features are important bases for identifying different symbols because symbols in military documents often have specific shapes and color codes to convey specific military information. Based on the extracted features, construct a shape analysis model that can identify and distinguish various symbol shapes. Subsequently, input the image data of non-text regions into the shape analysis model for processing. The model will judge which regions may be candidate regions containing important military symbols according to features such as the complexity of the symbol shape. Finally, combine detailed features such as line thickness and filling color to perform further annotation region matching for candidate symbol regions, thereby generating a set of candidate symbol annotation regions.
[0038] Step S400: For the set of candidate symbol annotation regions, use pattern recognition algorithms to match with a military symbol knowledge base and screen out real symbol annotation regions.
[0039] Specifically, for the set of candidate symbol annotation regions generated previously, use pattern recognition algorithms for detailed analysis. Pattern recognition algorithms can capture and analyze detailed features such as the complex shape features of symbols, line thickness, and filling color, which are important bases for distinguishing different military symbols. The system will compare these attributes of candidate symbol annotation regions with the standard codes and detailed group attributes in the military symbol knowledge base one by one. The military symbol knowledge base is a database containing a large number of military symbols and their related attributes, which provides key information such as the military meaning and usage scenarios of symbols. Through the close matching of pattern recognition algorithms with the military symbol knowledge base, the system can accurately identify which candidate symbol annotation regions are real military symbols, thereby screening out real symbol annotation regions.
[0040] Step S500: Determine the relevance between the real text regions and the real symbol annotation regions according to their relative positions, and construct a mapping relationship between text and symbol annotations.
[0041] Specifically, first, obtain the coordinate information of the real text area and the real symbol annotation area, which is the basis for determining their relative positions. Then, adopt a spatial relationship reasoning algorithm and use this coordinate information to calculate the distance value between the two. This calculation process takes into account the actual layout of the text and symbols in the document and can accurately reflect their spatial proximity. Next, compare the calculated distance value with a preset threshold, which is a key step in determining whether there is a correlation between the text and the symbol. If the distance value is less than the preset threshold, it indicates that the text and the symbol are close enough in space and are likely to have a logical or semantic correlation, so it is determined that there is a correlation between the two. Finally, based on this correlation, establish a mapping relationship between the real text area and the real symbol annotation area. This mapping relationship provides a basis for subsequent information structuring processing, ensuring that the text content and symbol annotations can maintain an accurate association in subsequent processing, thus helping to maintain the integrity and consistency of information and providing strong support for information extraction and knowledge graph construction of military documents.
[0042] Step S600: For the text area where the mapping relationship is established, use an OCR model to identify the text content and correct the recognition result according to the military term knowledge base.
[0043] Specifically, since military documents contain a large number of professional terms and specific formats, traditional OCR models cannot accurately identify these special contents. Therefore, the present invention further combines a military term knowledge base to correct the recognition result on the basis of OCR model recognition. When the OCR model outputs the preliminary recognition result, the system will compare these results with the standard terms in the military term knowledge base. If it is found that there are inconsistencies or errors between the recognition result and the terms in the knowledge base, the system will automatically correct the recognition result according to the correct terms in the knowledge base. This step ensures the accurate recognition of the text content, especially the accurate extraction of military terms, and provides a reliable data basis for subsequent information structuring and knowledge graph construction. By this method of combining the OCR model with the military term knowledge base, the present invention effectively improves the accuracy and efficiency of military document information extraction.
[0044] In a specific example, assume there is a PDF document containing a large number of military terms, and there is a passage of text describing the performance parameters of a new type of missile. In this step, the system first performs preliminary text content recognition on the text area for which the mapping relationship has been established (i.e., the text area describing the missile performance). However, due to the particularity of military terms, there may be errors in the preliminary recognition results of the OCR model. For example, "guidance accuracy" may be misrecognized as "system accuracy". Subsequently, the system corrects the recognition results according to the military term knowledge base. The knowledge base stores the standard term "guidance accuracy" and its related explanations. By comparing the recognition results with the content of the knowledge base, the system automatically corrects the incorrect "system accuracy" to the correct "guidance accuracy". This step not only improves the accuracy of text content recognition but also ensures the reliability of subsequent information structuring processing and knowledge graph construction. Among them, the construction method of the OCR model can be, for example: First, for the text area in the military document, extract its image data and perform preprocessing, including steps such as grayscale conversion and binarization to improve the image quality. Then, use deep learning algorithms (such as convolutional neural network CNN) to perform feature extraction and pattern recognition training on the preprocessed image, enabling the model to learn features such as the shape and structure of text characters. During the training process, a large number of labeled military document images are used as training samples to improve the model's recognition ability for military terms and specific formats. Finally, by combining the military term knowledge base, post-process the recognition results of the OCR model to correct misrecognized or missing military terms, thereby constructing an OCR model that can accurately recognize the text content in military documents.
[0045] Step S700, for the symbol annotation area with the mapping relationship established, obtain symbol annotation information using a symbol recognition model, and correct the recognition result according to the military symbol knowledge base;
[0046] Specifically, the symbol recognition model first deeply analyzes the image data of the symbol annotation area, extracts key visual features such as the shape, lines, and filling color of the symbol. Subsequently, the model uses these features for preliminary recognition and outputs the annotation information of the symbol. However, the preliminary recognition results may not be completely accurate because military symbols often have specific meanings and usages, which require combining professional knowledge in the military field to understand. Therefore, the system further corrects the recognition results according to the military symbol knowledge base. The military symbol knowledge base stores information such as the standard codes, detail group attributes, military meanings, and usage scenarios of a large number of military symbols. By comparing this information with the recognition results, the system can verify the accuracy of the recognition results and correct possible errors.
[0047] In a specific example, taking a military document containing a tactical map as an example, there are multiple symbol annotations on the map indicating the positions of different military units, such as "infantry unit symbol", "artillery unit symbol", etc. These symbol annotations are crucial for understanding the tactical layout on the map. The system first identifies these symbol annotation areas through the previous steps and establishes the mapping relationship between them and the text areas. Then, in this step, the symbol recognition model identifies each symbol annotation area and corrects the recognition results according to the military symbol knowledge base. For example, for the "infantry unit symbol", the model may initially identify its shape and color, but the knowledge base will further confirm its specific military meaning and usage scenario to ensure the accuracy of the recognition result. Finally, these corrected symbol annotation information, together with the text content, will be used to generate structured military element description data and construct a complete military element knowledge graph. Among them, the symbol recognition model can be constructed, for example, through the following steps: First, obtain the image data of the non-text area, and use the feature extraction algorithm to extract key features such as the shape complexity, line thickness, and filling color of the symbol; then, based on these features, construct a shape analysis model, input the non-text area into the model for preliminary screening, and generate a set of candidate symbol annotation areas; next, extract the standard code and detail group attributes of the candidate symbol annotation areas, match them with the military meaning and usage field attributes in the military symbol knowledge base, and use the pattern recognition algorithm to further verify the authenticity of the candidate symbol areas, and screen out the real symbol annotation areas; finally, according to the characteristics of the real symbol annotation areas, train and optimize the symbol recognition model so that it can accurately identify and extract the symbol annotation information in the military document, and correct the recognition results in combination with the military symbol knowledge base to ensure the accuracy and reliability of symbol recognition.
[0048] Step S800: According to the mapping relationship between the text and the symbol annotations, associate the corrected text recognition results and symbol annotation information to generate structured military element description data and construct a military element knowledge graph;
[0049] Specifically, accurate text content and symbol annotation information corrected by the military term knowledge base and the military symbol knowledge base are first obtained, which serve as the basis for constructing structured military element description data. Subsequently, using structured processing techniques, this information is organized according to a predetermined data structure and format to generate military element description data that is easy to analyze and retrieve. This process not only ensures the accuracy and consistency of the data but also improves its usability. In addition, in this step, through a knowledge reasoning algorithm, the generated military element description data is compared with relevant information in the military knowledge base to further enrich and improve the description content of military elements. The comparison results not only verify the accuracy of the identification information but may also reveal potential connections and patterns between military elements. Finally, based on the comparison results, the effectiveness attributes of military elements are inferred and added to the knowledge graph, thereby constructing a military element knowledge graph that contains rich military element information, has a clear structure, is easy to analyze, and supports decision-making. This step not only realizes the effective extraction and structured processing of military document information but also provides strong data support and knowledge foundation for information retrieval, analysis, and decision-making in the military field.
[0050] In this embodiment, step S100 includes:
[0051] Step S101, extracting text line spacing, alignment, font size, and color information from the military document PDF file;
[0052] Specifically, in this step, by parsing the structure of the PDF file, detailed information related to the text layout is extracted. The line spacing refers to the vertical distance between two lines of text, which is of great significance for determining whether the text belongs to the same paragraph or the same area. The alignment reflects the arrangement of the text on the page, such as left alignment, right alignment, center alignment, or justified alignment, and this information helps to identify the organizational structure and logical relationship of the text. The font size determines the visual presentation effect of the text, and different font sizes may represent different heading levels or emphasis degrees. The color information is used to distinguish the type or purpose of the text. For example, red or bold text may indicate important information or warnings. Suppose there is a military document containing multiple paragraphs and headings. In this step, the system will first analyze the line spacing of the document and find that the distance between certain lines is significantly smaller, indicating that these lines belong to the same paragraph. Then, the system will check the alignment and find that most of the text is left-aligned, but a part of the text is center-aligned, and this part of the text may be a heading or an important statement. In addition, the system will also extract the font size and color information and find that the headings are in a larger and bold black font, while the body text is in a standard-size black font, and some key terms or data are highlighted in red font.
[0053] Step S102, according to the preset rule algorithm, if the line spacing of the text is less than the preset threshold, it is determined to be the same text area;
[0054] Specifically, the system presets a threshold for line spacing. When processing a PDF document, it analyzes the line spacing of the text lines in the document. If the line spacing between two paragraphs of text is less than this preset threshold, the system determines that these two paragraphs of text belong to the same text area. This judgment is based on the assumption that text that is typeset closer together is more likely to belong to the same logical or semantic unit. Take an example to illustrate. Suppose there is a military report document that contains a chapter title and several paragraphs of text immediately following it. In the PDF, the line spacing between the chapter title and the first paragraph of text may be very small, such as only 10 pixels, while the line spacing between paragraphs may be 20 pixels. If the preset line spacing threshold is 15 pixels, then when processing this document, the system will determine that the chapter title and the first paragraph of text belong to the same text area because their line spacing is less than 15 pixels. On the contrary, the line spacing between the first paragraph of text and the second paragraph of text is greater than 15 pixels, and the system will regard them as two independent text areas (although they both belong to the same chapter). This step can help the system divide the text area more accurately, thereby providing accurate basic data for subsequent text recognition, information structuring, and knowledge graph construction. If the text area is divided inaccurately, for example, text from different chapters or different topics is wrongly grouped together, then subsequent information extraction and knowledge graph construction will be incorrect, resulting in unreliable final analysis results.
[0055] Step S103, if the difference in font size exceeds the preset range, it is divided into different text areas;
[0056] Specifically, when processing military documents, since the document may contain various levels of text information such as titles, text, and annotations, the font sizes of these texts often vary. For example, titles usually use a larger font size to highlight their importance, while the text uses a relatively smaller font size to maintain reading comfort. If these different levels of text are wrongly classified into the same text area, it may lead to confusion in subsequent information extraction and structuring processes. Therefore, in this step, by comparing the difference in font size between texts, when the difference exceeds the preset range, these texts are divided into different text areas. The advantage of doing this is that it can more accurately reflect the structure and content hierarchy of the document, providing clearer and more accurate basic data for subsequent information extraction and knowledge graph construction.
[0057] Step S104, according to the font color feature, the text content with the same color is grouped into the same candidate area;
[0058] Specifically, military documents may contain text in multiple colors, which are often used to distinguish different types of information, such as headings, body text, annotations, or emphasized content. By identifying and analyzing these font color features, the system can group text content with the same color into the same candidate region. Suppose a military document contains red headings, black body text, and blue annotations. In this step, the system will first identify all the red, black, and blue text content in the document. Then, the system will classify the text with the same color into different candidate regions respectively. For example, all red text will be classified into a "red text candidate region", all black text will be classified into a "black text candidate region", and all blue text will be classified into a "blue text candidate region".
[0059] Step S105, using a rule-based region merging algorithm, merge adjacent candidate text regions with the same alignment;
[0060] Specifically, in military documents, due to the complexity of typesetting and formatting, the initially divided text regions may be too fragmented, which is not conducive to subsequent text recognition and information extraction. Therefore, through the region merging algorithm, text regions that visually belong to the same logical block can be merged, thereby improving the accuracy and efficiency of text recognition. Suppose there is a table in a military document, and a certain column in the table contains multiple text items with the same alignment, such as "equipment name", "quantity", etc. In the initial region division, these text items may be divided into multiple independent candidate text regions. However, logically, they belong to the same column and should be treated as a whole. Therefore, in this step, the region merging algorithm will identify these adjacent candidate text regions with the same alignment and merge them into a larger text region. In this way, subsequent text recognition and information extraction can be performed on this larger text region, thereby improving the accuracy and efficiency of processing. In addition, the region merging algorithm can also be merged according to other rules, such as font size, color and other features. These rules can be customized and adjusted according to specific military documents and recognition requirements.
[0061] Step S106, through a preset layout feature combination rule, eliminate the regions that do not meet the requirements to generate the set of candidate text regions;
[0062] Specifically,
[0063] In this embodiment, the step S200 includes:
[0064] Step S201, extract the text content in the set of candidate text regions to form a text feature set;
[0065] Step S202: According to the military term knowledge base, obtain the standard name, synonyms, hyponymy / hypernymy relationships, and English translation attributes, and construct term matching rules;
[0066] Step S203: Adopt a term matching algorithm to match the text feature set with the term matching rules one by one, and calculate the matching degree;
[0067] Step S204: If the matching degree is higher than the preset threshold, it is determined as a real text area;
[0068] Step S205: According to the matching results, screen out all the real text areas;
[0069] Specifically, military documents often contain a large number of professional terms and content in a specific format, which are crucial for understanding the overall meaning and details of the documents. Traditional OCR technology cannot accurately recognize these professional terms because they may not be in a general dictionary or language model. Therefore, the present invention introduces a military term knowledge base, which contains rich information such as standard terms, synonyms, hyponymy / hypernymy relationships, and English translations in the military field. Specifically, first, the system extracts the text content from the set of candidate text areas to form a text feature set. These text features may include, for example, language features such as vocabulary, phrases, and sentence structures, as well as visual features such as fonts and layouts. Then, the system constructs term matching rules according to the military term knowledge base. These rules may be defined based on attributes such as the standard name, synonyms, hyponymy / hypernymy relationships, and English translations of terms. For example, if a term has multiple synonyms in the knowledge base, the matching rules will take these synonyms into account. Then, the system adopts a term matching algorithm to match the text feature set with the term matching rules one by one. In this process, the algorithm calculates the similarity or matching degree between the text features and the matching rules. If the matching degree of a text feature with a certain term is higher than the preset threshold, then the system will determine that this text area contains real and important military information, that is, it is determined as a real text area. Finally, the system will screen out all real text areas according to the matching results. These areas will serve as the basis for subsequent information structuring processing and knowledge graph construction. In this way, the present invention can ensure that the text content extracted from military documents is accurate, complete, and relevant to the military field, thereby improving the accuracy and efficiency of information extraction.
[0070] In a specific example, assume that the military term knowledge base contains a term "Air-to-Air Missile", and its synonyms include "Air-Air Missile". The text content of candidate text area R 1 is "Air-to-Air Guide", and the text content of candidate text area R 2 is "Surface-to-Air Missile". For R 1 , calculate the matching degree: (Because there is only one term "air-to-air" related to the term, but "guide" does not match). For R 2 , calculate the matching degree: MatchDegree(T 2 , "air-to-air missile") = 0 (because no terms are related to the term). Assume the preset threshold θ = 0.2, then R 1 will be determined as a real text area (although not a perfect match, the matching degree is higher than the threshold), while R 2 will not be determined as a real text area.
[0071] In this embodiment, the step S300 includes:
[0072] Step S301, obtaining the image data of the non-text area;
[0073] Step S302, using a feature extraction algorithm to obtain the symbol shape complexity, line thickness, and filling color from the image data;
[0074] Step S303, constructing a shape analysis model according to the symbol shape features;
[0075] Step S304, inputting the non-text area into the shape analysis model for processing. If the symbol shape complexity is higher than the preset threshold, it is determined as a candidate symbol area;
[0076] Step S305, performing annotation area matching on the candidate symbol area according to the line thickness and the filling color to generate the set of candidate symbol annotation areas;
[0077] Specifically, first, the system extracts the image data of non-text areas from the PDF file of military documents. These non-text areas may contain various complex military symbols or other graphic elements. To accurately identify these symbols, the system uses a feature extraction algorithm to analyze the image data and extract key features such as the shape complexity, line thickness, and filling color of the symbols. Shape complexity is an important recognition feature because military symbols often have specific shapes and structures, which are usually more complex than ordinary text or graphics. By calculating the shape complexity of the symbols, the system can initially screen out the areas that may be military symbols. Line thickness and filling color are also important features for identifying symbols because different symbols may have different line thicknesses and filling colors, and these features help to further distinguish and identify the symbols. Next, the system constructs a shape analysis model based on the extracted symbol shape features. This model can identify and analyze the shape features of various symbols and classify and determine the symbols according to preset rules and algorithms. When the image data of the non-text area is input into the shape analysis model, the model will analyze and process each area in the image to determine whether it conforms to the characteristics of military symbols. If the shape complexity of the symbol in a certain area is higher than the preset threshold, the system will determine it as a candidate symbol area. This means that this area may contain one or more military symbols and requires further analysis and confirmation. Finally, the system performs annotation area matching on the candidate symbol area according to features such as line thickness and filling color. This step is to further verify and confirm the accuracy of the candidate symbol area. The system will compare the features of the candidate symbol area with the standard symbols in the military symbol knowledge base, find the most matching symbol, and generate a set of candidate symbol annotation areas.
[0078] In a specific example, for instance, a military document contains a symbol representing "air raid", which consists of a complex graphic and a specific filling color. After extracting the image data of the non-text area, the system obtains features such as the shape complexity, line thickness, and filling color of this symbol through the feature extraction algorithm. Then, the system inputs these features into the shape analysis model for processing. Since the shape complexity of this symbol is relatively high and it conforms to the characteristics of military symbols, the system determines it as a candidate symbol area. Finally, the system finds the most matching symbol in the military symbol knowledge base according to features such as line thickness and filling color and adds it to the set of candidate symbol annotation areas. Through this process, the system can accurately identify the symbol annotation areas in military documents, providing an accurate data basis for subsequent information extraction and knowledge graph construction. This method based on shape analysis not only improves the accuracy of symbol recognition but also enhances the adaptability of the system to military documents in different formats and layouts.
[0079] In this embodiment, step S400 includes:
[0080] Step S401, extract the standard code and detail group attributes from the set of candidate symbol annotation regions;
[0081] Step S402, according to the standard code and the detail group attributes, obtain the military meaning and usage field attributes from the military symbol knowledge base;
[0082] Step S403, adopt a pattern recognition algorithm to compare the attributes of shape complexity, line thickness, and filling color in the candidate symbol region with a preset threshold;
[0083] Step S404, if the matching value output by the pattern recognition algorithm is higher than the preset threshold, determine that the candidate symbol region is the real symbol annotation region;
[0084] Specifically, the standard code is, for example, the unique identifier of the symbol, and the detail group attributes may include, for example, visual characteristics such as the shape feature description of the symbol, line features, and filling color. These attributes are the basis for identifying and understanding military symbols because they are usually designed with specific visual features to convey specific military information. The system accesses the military symbol knowledge base according to the extracted standard code and detail group attributes. This knowledge base is a database containing a large number of military symbols and their related attributes, which provides key information such as the military meaning (i.e., the meaning of the symbol in the military context) and usage field (i.e., the scenarios or conditions applicable to the symbol) of each symbol. By matching these attributes with the information in the knowledge base, the system can obtain the detailed explanations and usages of the symbols, which are important bases for judging the authenticity and accuracy of the symbols. Subsequently, the system uses a pattern recognition algorithm to further analyze the candidate symbol region. This algorithm compares the attributes such as the shape complexity, line thickness, and filling color of the candidate symbol with the preset thresholds. These thresholds are set based on the standard design and usage rules of military symbols, and they help the system distinguish real symbols from possible misidentifications or noises. Finally, if the matching value output by the pattern recognition algorithm is higher than the preset threshold, the system will determine that the candidate symbol region is the real symbol annotation region. This means that the system believes that the symbol contained in this region is the symbol actually used in the military document with a specific military meaning, rather than a misidentification caused by image distortion, noise, or other factors.
[0085] In a specific example, a military document contains a symbol representing "air force base", and this symbol has corresponding standard codes and detailed group attributes (such as specific shape, line thickness, and filling color) in the military symbol knowledge base. When the system processes this document, it will first identify this symbol and extract its attributes. Then, the system will access the military symbol knowledge base to find the standard codes and attributes that match this symbol, thereby confirming that this symbol represents "air force base". Next, the system will adopt a pattern recognition algorithm to compare the attributes of this symbol with a preset threshold. If the matching degree is high, it will determine that this symbol area is real and process it as part of military elements. For example, the standard shape of the symbol "air force base" is a specific geometric figure, the line thickness is 2 pixels, and the filling color is blue. The candidate symbol annotation area S 1 has a shape similar to the "air force base" symbol, the line thickness is 2.1 pixels, and the filling color is dark blue (with a small Euclidean distance from the standard blue in the RGB space). Shape similarity: For example, the shape similarity calculated by the Hausdorff distance is 0.9 (indicating very similar). Line thickness similarity: The absolute difference is ∣2.1 - 2∣ = 0.1, which can be normalized to 0.1 / 2 = 0.05 (indicating very close). Filling color similarity: For example, the color similarity calculated by the Euclidean distance in the RGB space is 0.95 (indicating very close in color). Assume the weight coefficients α = 0.5, β = 0.3, γ = 0.2, then the matching degree is:
[0086] θ = MatchDegree(S 1 , "air force base") = 0.5·0.9 + 0.3·0.05 + 0.2·0.95
[0087] = 0.45 + 0.015 + 0.19 = 0.655
[0088] Assume the preset threshold θ = 0.6, then MatchDegree(S 1 , "air force base") > θ, so S 1 is determined to be a real symbol annotation area.
[0089] Based on the above analysis, it can be seen that the present invention can accurately identify and extract symbol annotations in military documents. Even if these symbols are closely adjacent to or overlapping with text in a complex layout, they can be effectively distinguished, which greatly improves the accuracy and efficiency of military document information processing.
[0090] In this embodiment, the step S500 includes:
[0091] Step S501, obtaining the coordinate information of the real text area and the real symbol annotation area, and calculating the relative position between the two; d
[0092] Step S502: Using a spatial relationship reasoning algorithm, calculate the distance value between the real text region and the real symbol annotation region according to the coordinate information;
[0093] Step S503: Compare the calculated distance value with a preset threshold. If the distance value is less than the preset threshold, it is determined that there is a correlation between the two;
[0094] Step S504: Establish a mapping relationship between the real text region and the real symbol annotation region to generate the mapping relationship between the text and the symbol annotation;
[0095] Specifically, the coordinate information refers to the specific positions of these regions on the document page, usually represented by the coordinates of the upper left corner and the width and height of the region. After obtaining this coordinate information, the system can calculate the relative positions between the text region and the symbol annotation region, such as the horizontal distance and the vertical distance between them. Then, the system uses a spatial relationship reasoning algorithm to process this coordinate information. The spatial relationship reasoning algorithm is an algorithm that can understand and analyze the relationships between objects in space. It can calculate the distance between objects according to their coordinate information and judge whether there is a spatial correlation between these objects according to preset rules. In the present invention, the algorithm calculates the distance value between the text region and the symbol annotation region, and this distance value reflects their proximity on the document page. Then, the system compares the calculated distance value with a preset threshold. This threshold is a preset value used to determine whether the text region and the symbol annotation region are close enough to possibly have a correlation. If the calculated distance value is less than this threshold, the system determines that these two regions are correlated. This judgment of correlation is based on spatial proximity, that is, it is considered that adjacent text and symbols in the document are likely to be related to each other. Finally, the system establishes a mapping relationship between the real text region and the real symbol annotation region. This mapping relationship is a data structure used to represent the association between the text region and the symbol annotation region. By establishing this mapping relationship, the system can accurately associate the text content and the symbol annotation information, providing a basis for subsequent information structuring processing and knowledge graph construction.
[0096] In a specific example, in a military document, there is a text area describing "artillery troops", and there is a symbol annotation area representing the artillery troops beside it (such as an icon of a gun barrel). The system first obtains the coordinate information of these two areas, and then calculates the distance value between them. If this distance value is less than a preset threshold, the system determines that there is a correlation between these two areas and establishes a mapping relationship between them. In this way, in subsequent information structuring processing, the system can accurately associate the text content of "artillery troops" with the icon of the gun barrel, generate structured military element description data, and further construct a military element knowledge graph. The establishment of this correlation is very important for understanding and analyzing the information in military documents. It helps to maintain the integrity and consistency of information and improve the accuracy and efficiency of information extraction. For example, there are the following two areas:
[0097] Text area: upper left corner coordinates (100, 200), width 200, height 50; symbol annotation area: upper left corner coordinates (130, 220), width 50, height 30; preset threshold T = 50, calculate the center point coordinates:
[0098] Center point of the text area:
[0099] Center point of the symbol annotation area:
[0100] Calculate the distance value:
[0101]
[0102] Judge the correlation: Since d = 46.10 < 50 = T, it is determined that there is a correlation between these two areas.
[0103] Construct the mapping relationship: Assume that the identifier of the text area is "text_region_1", and the identifier of the symbol annotation area is "symbol_region_1".
[0104] The mapping relationship can be expressed as: {"text_region_1": "symbol_region_1"}.
[0105] In this embodiment, the step S800 includes:
[0106] Step S801, obtain the corrected text recognition result and symbol annotation information;
[0107] Step S802, adopt structuring processing technology to generate military element description data;
[0108] Step S803: Use a knowledge inference algorithm to extract relevant information from the military knowledge base based on the military element description data for comparison and generate a comparison result;
[0109] Step S804: Infer the effectiveness attributes of military elements based on the comparison result;
[0110] Step S805: Add the effectiveness attributes of the military elements to the knowledge graph to generate the military element knowledge graph;
[0111] Specifically, when the system obtains preliminary text and symbol recognition results through the OCR model and the symbol recognition model, there may be certain errors or inaccuracies in these results. Therefore, the system first corrects these recognition results using the military term knowledge base and the military symbol knowledge base to ensure the accuracy of the text content and symbol annotations. The system adopts structured processing technology to organize the corrected text recognition results and symbol annotation information according to a predetermined data structure and format to generate military element description data. This process not only includes associating text and symbol information but may also involve classifying, summarizing, and organizing the information to form a clearer, more understandable, and analyzable data structure. Subsequently, the system uses a knowledge inference algorithm to compare the generated military element description data with relevant information in the military knowledge base. The purpose of this step is to further verify the accuracy of the recognition information and reveal the potential connections and rules between military elements. For example, a certain military term may have a hyponymy or hypernymy relationship with other terms, or a certain symbol may have different meanings and usages in different scenarios. Through the knowledge inference algorithm, the system can discover these potential connections and accordingly enrich and improve the description content of military elements. Based on the comparison result, the system further infers the effectiveness attributes of military elements. The effectiveness attribute refers to the role, effect, or value of a military element in a specific scenario. For example, the range, accuracy, and lethality of a certain weapon system are its effectiveness attributes. By inferring the effectiveness attributes, the system can more comprehensively understand the characteristics and performance of military elements, providing strong support for subsequent decision-making and analysis. Finally, the system adds the effectiveness attributes of military elements to the knowledge graph to generate the military element knowledge graph. The knowledge graph is a structured knowledge representation method that can present entities, attributes, and relationships in a graphical way for easy understanding and analysis. By constructing the military element knowledge graph, the system can display the extracted military element information in a more intuitive and clear manner, providing strong data support and knowledge foundation for information retrieval, analysis, and decision-making in the military field.
[0112] In a specific implementation, the system extracts a description of a new type of tank from a military document, including text content such as its model, weight, and firepower configuration, as well as related symbolic annotations (such as symbols representing the tank's firepower). The system first corrects this information using a military terminology knowledge base and a military symbol knowledge base to ensure its accuracy. Then, the system uses structured processing techniques to organize this information into structured military element description data, such as "new type of tank model X, weight Y tons, equipped with Z-type guns", etc. Next, the system compares this description data with relevant information in the military knowledge base through a knowledge reasoning algorithm and finds that the firepower configuration of this tank is similar to that of a certain known tank, but it is lighter in weight and may therefore have better mobility. Based on this comparison result, the system infers that the effectiveness attributes of this tank may include "higher mobility and firepower", and adds these attributes to the knowledge graph to generate a knowledge graph about this new type of tank. In this way, users can quickly understand the main features and performance of this tank by browsing the knowledge graph, providing strong support for subsequent decision-making and analysis.
[0113] Specifically, for example, cosine similarity can be used to calculate the attribute matching degree, and the formula is as follows:
[0114]
[0115] Where: is the attribute vector of the military element description data, is the attribute vector of the matching entry in the military knowledge base; is the dot product of the vectors and ; and are the magnitudes of the vectors and
[0116] If the calculated cosine similarity is higher than a preset threshold (such as 0.8), it is considered that the attribute matching is successful, and the effectiveness attributes are extracted from the matching entry for inference.
[0117] Taking a text describing a new type of tank in a military document as an example, after OCR recognition and correction by a military term knowledge base, the following military element description data is obtained: Model: T-90; Weight: 46 tons; Firepower configuration: 125 mm smoothbore gun. Search for entries in the military knowledge base that match these attributes. Suppose a highly matching entry is found, and its effectiveness attributes include: Firepower coverage: 3000 meters; Mobility: High (maximum speed 65 km / h); Concealment: Medium (thicker armor, but obvious infrared characteristics). According to the specific situation in the document (such as the combat environment being a plain area), the weights or values of the effectiveness attributes can be adjusted. For example, increase the weight of mobility because higher mobility is required in the plain area. Finally, infer the effectiveness attributes of the new type of tank, generate knowledge graph nodes containing these attributes, and add them to the knowledge graph.
[0118] Example two, Figure 2 is a schematic structural diagram of a military document information extraction and knowledge graph construction system in the second embodiment of the present invention, as Figure 2As shown in the figure, Embodiment 2 provides a military document information extraction and knowledge graph construction system, including: a candidate text region generation module 201, a real text region screening module 202, a candidate symbol annotation region generation module 203, a real symbol annotation region screening module 204, a mapping relationship construction module 205, a text region correction module 206, a symbol annotation region correction module 207, and a knowledge graph construction module 208. The candidate text region generation module 201 is used to obtain a military document PDF file, divide text regions according to text layout features, and generate a set of candidate text regions. The real text region screening module 202 is used to match the set of candidate text regions with a military term knowledge base by using a term matching algorithm, and screen out real text regions. The candidate symbol annotation region generation module 203 is used to perform shape analysis on non-text regions according to symbol shape features, and generate a set of candidate symbol annotation regions. The real symbol annotation region screening module 204 is used to match the set of candidate symbol annotation regions with a military symbol knowledge base by using a pattern recognition algorithm, and screen out real symbol annotation regions. The mapping relationship construction module 205 is used to determine the relevance between the real text region and the real symbol annotation region according to their relative positions, and construct a mapping relationship between the text and the symbol annotation. The text region correction module 206 is used to recognize the text content of the text region with a mapping relationship established by using an OCR model, and correct the recognition result according to the military term knowledge base. The symbol annotation region correction module 207 is used to obtain symbol annotation information for the symbol annotation region with a mapping relationship established by using a symbol recognition model, and correct the recognition result according to the military symbol knowledge base. The knowledge graph construction module 208 is used to associate the corrected text recognition result and symbol annotation information according to the mapping relationship between the text and the symbol annotation, generate structured military element description data, and construct a military element knowledge graph.
[0119] In this embodiment, the candidate text region generation module 201 includes: an information extraction unit, a region determination unit, a region division unit, a region classification unit, a region merging unit, and a set generation unit. The information extraction unit is used to extract text line spacing, alignment, font size, and color information from the military document PDF file. The region determination unit is used to determine the same text region according to a preset rule algorithm if the text line spacing is less than a preset threshold. The region division unit is used to divide into different text regions if the font size difference exceeds a preset range. The region classification unit is used to classify text content with the same color into the same candidate region according to the font color feature. The region merging unit is used to merge adjacent candidate text regions with the same alignment by using a rule-based region merging algorithm. The set generation unit is used to generate the candidate text region set by eliminating non-conforming regions through a preset layout feature combination rule.
[0120] In this embodiment, the real text region screening module 202 includes: a feature set formation unit, a matching rule construction unit, a matching degree calculation unit, a text region determination unit, and a text region screening unit. The feature set formation unit is used to extract the text content in the candidate text region set to form a text feature set. The matching rule construction unit is used to obtain standard names, synonyms, hyponymy and hypernymy relationships, and English translation attributes according to the military term knowledge base to construct term matching rules. The matching degree calculation unit is used to use a term matching algorithm to match the text feature set with the term matching rules item by item and calculate the matching degree. The text region determination unit is used to determine a real text region if the matching degree is higher than a preset threshold. The text region screening unit is used to screen out all the real text regions according to the matching results.
[0121] In this embodiment, the candidate symbol annotation region generation module 203 includes: an image data acquisition unit, a first acquisition unit, an analysis model construction unit, a symbol region determination unit, and an annotation region set generation unit. The image data acquisition unit is used to acquire the image data of the non-text region. The first acquisition unit is used to acquire symbol shape complexity, line thickness, and filling color from the image data by using a feature extraction algorithm. The analysis model construction unit is used to construct a shape analysis model according to the symbol shape feature. The symbol region determination unit is used to input the non-text region into the shape analysis model for processing, and determine it as a candidate symbol region if the symbol shape complexity is higher than a preset threshold. The annotation region set generation unit is used to perform annotation region matching on the candidate symbol region according to the line thickness and the filling color to generate the candidate symbol annotation region set.
[0122] In this embodiment, the module 204 for screening real symbol annotation regions includes: an extraction unit, a second acquisition unit, a comparison unit, and a region determination unit. The extraction unit is configured to extract the standard code and the detail group attributes from the set of candidate symbol annotation regions. The second acquisition unit is configured to acquire the military meaning and usage field attributes from the military symbol knowledge base according to the standard code and the detail group attributes. The comparison unit is configured to compare the attributes of the shape complexity, line thickness, and filling color in the candidate symbol region with a preset threshold by using a pattern recognition algorithm. The region determination unit is configured to determine that the candidate symbol region is the real symbol annotation region if the matching value output by the pattern recognition algorithm is higher than the preset threshold.
[0123] In this embodiment, the module 205 for constructing a mapping relationship includes: a position calculation unit, a distance value calculation unit, a relevance determination unit, and a mapping relationship generation unit. The position calculation unit is configured to acquire the coordinate information of the real text region and the real symbol annotation region, and calculate the relative position between the two. The distance value calculation unit is configured to calculate the distance value between the real text region and the real symbol annotation region according to the coordinate information by using a spatial relationship reasoning algorithm. The relevance determination unit is configured to compare the calculated distance value with a preset threshold, and determine that there is a relevance between the two if the distance value is less than the preset threshold. The mapping relationship generation unit is configured to establish a mapping relationship between the real text region and the real symbol annotation region, and generate the mapping relationship between the text and the symbol annotation.
[0124] In this embodiment, the module 208 for constructing a knowledge graph includes: a third acquisition unit, a description data generation unit, a comparison result generation unit, an effectiveness attribute inference unit, and a knowledge graph generation unit. The third acquisition unit is configured to acquire the corrected text recognition result and symbol annotation information. The description data generation unit is configured to generate military element description data by using a structured processing technology. The comparison result generation unit is configured to extract relevant information from the military knowledge base based on the military element description data through a knowledge reasoning algorithm for comparison, and generate a comparison result. The effectiveness attribute inference unit is configured to infer the effectiveness attribute of the military element based on the comparison result. The knowledge graph generation unit is configured to add the effectiveness attribute of the military element to the knowledge graph, and generate the knowledge graph of the military element.
[0125] The various variations and specific examples of the method for extracting military document information and constructing a knowledge graph provided in the first embodiment are equally applicable to the system for extracting military document information and constructing a knowledge graph provided in this embodiment. Through the foregoing detailed description of a method for extracting military document information and constructing a knowledge graph, those skilled in the art can clearly know the implementation manner of the system for extracting military document information and constructing a knowledge graph in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated here.
[0126] Embodiment 3 Figure 3 is a schematic structural diagram of an electronic device in Embodiment 3 of the present invention. As Figure 3 shown, Embodiment 3 further provides an electronic device 300, and the electronic device may include: a processor 301 and a memory 302.
[0127] The memory 302 is used to store programs; the memory 302 may include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviation: RAM), such as static random access memory (English: static random-access memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory 302 is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. may be stored in one or more memories 302 in a partitioned manner. And the above computer programs, computer instructions, data, etc. may be called by the processor 301.
[0128] The above computer programs, computer instructions, etc. may be stored in one or more memories 302 in a partitioned manner. And the above computer programs, computer data, etc. may be called by the processor 301. The processor 301 is used to execute the computer programs stored in the memory 302 to implement each step in the methods involved in the above embodiments. For specific reference, see the relevant descriptions in the foregoing method embodiments.
[0129] The processor 301 and the memory 302 can be independent structures or integrated structures integrated together. When the processor 301 and the memory 302 are independent structures, the memory 302 and the processor 301 can be coupled through the bus 303. The electronic device of this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0130] Embodiment 4 also provides a computer-readable storage medium, including a computer program and instructions. When the computer program or instructions run on a computer, the computer is enabled to execute the military document information extraction and knowledge graph construction method of any embodiment of the present invention.
[0131] The computer-readable storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0132] This embodiment also provides a computer program product. The computer program product includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to enable the electronic device to execute the solution provided in any of the above embodiments.
[0133] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recorded in the disclosure of the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. No limitations are imposed herein.
[0134] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for extracting military document information and constructing a knowledge graph, characterized in that: include: Step S100, obtaining a military document PDF file, dividing the text area according to the text layout features, and generating a candidate text area set; Step S200, for the candidate text region set, a term matching algorithm is used to match with a military term knowledge base to screen out the real text region; Step S300, for the non-text area, shape analysis is performed according to the symbol shape features to generate a set of candidate symbol annotation areas; Step S400, for the candidate symbol annotation area set, a pattern recognition algorithm is used to match with a military symbol knowledge base to screen out the real symbol annotation area; Step S500, determining the correlation between the real text area and the real symbol annotation area according to their relative positions, and constructing a mapping relationship between the text and the symbol annotation; Step S600, for the text area with established mapping relationship, using OCR model to recognize text content, and correcting the recognition result according to the military terminology knowledge base; Step S700, for the symbol annotation area with established mapping relationship, using the symbol recognition model to obtain symbol annotation information, and correcting the recognition result according to the military symbol knowledge base; Step S800, according to the mapping relationship between the text and the symbol annotation, the corrected text recognition result and the symbol annotation information are associated to generate structured military element description data and construct a military element knowledge graph.
2. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: The obtaining of a military document PDF file, dividing the text area according to the text layout features, and generating a candidate text area set comprises: extracting text line spacing, alignment, font size and color information from the military document PDF file; According to the preset rule algorithm, if the text line spacing is less than the preset threshold, it is determined to be the same text area; If the font size difference exceeds the preset range, it is divided into different text areas; According to the font color features, the text contents of the same color are classified into the same candidate area; A rule-based region merging algorithm is used to merge adjacent candidate text regions with the same alignment. The candidate text region set is generated by eliminating regions that do not meet the requirements through preset layout feature combination rules.
3. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: The candidate text region set is matched with the military terminology knowledge base using a term matching algorithm to screen out the real text regions, including: Extracting text content from the candidate text region set to form a text feature set; According to the military terminology knowledge base, standard names, synonyms, hyponyms and English translation attributes are obtained to construct term matching rules; Using a term matching algorithm, the text feature set is matched with the term matching rule one by one, and a matching degree is calculated; If the matching degree is higher than the preset threshold, it is determined to be a real text area; According to the matching results, all the real text regions are screened out.
4. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: The step of performing shape analysis on the non-text area according to the symbol shape features to generate a set of candidate symbol annotation areas includes: Acquiring image data of the non-text area; Using a feature extraction algorithm to obtain symbol shape complexity, line thickness and fill color from the image data; According to the shape characteristics of the symbol, a shape analysis model is constructed; The non-text area is input into the shape analysis model for processing, and if the symbol shape complexity is higher than a preset threshold, it is determined to be a candidate symbol area; According to the line thickness and the filling color, annotation area matching is performed on the candidate symbol area to generate the candidate symbol annotation area set.
5. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: The candidate symbol annotation area set is matched with the military symbol knowledge base by using a pattern recognition algorithm to screen out the real symbol annotation areas, including: extracting standard codes and detail group attributes in the candidate symbol annotation region set; Acquire military meaning and usage field attributes from a military symbol knowledge base according to the standard code and the detail group attributes; Using a pattern recognition algorithm, the attributes of shape complexity, line thickness, and fill color in the candidate symbol area are compared with preset thresholds; If the matching value output by the pattern recognition algorithm is higher than a preset threshold, the candidate symbol region is determined to be the real symbol annotation region.
6. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: Determining the correlation between the real text area and the real symbol annotation area according to their relative positions, and constructing a mapping relationship between the text and the symbol annotation includes: Obtaining coordinate information of the real text area and the real symbol annotation area, and calculating the relative position between the two; Using a spatial relationship reasoning algorithm, the distance value between the real text area and the real symbol annotation area is calculated according to the coordinate information; The calculated distance value is compared with a preset threshold value. If the distance value is less than the preset threshold value, it is determined that there is a correlation between the two. A mapping relationship between the real text area and the real symbol annotation area is established to generate a mapping relationship between the text and the symbol annotation.
7. The method for extracting military document information and constructing a knowledge graph as claimed in claim 1, characterized in that: According to the mapping relationship between text and symbol annotation, the corrected text recognition result and symbol annotation information are associated to generate structured military element description data, and the military element knowledge graph is constructed, including: Obtain corrected text recognition results and symbol annotation information; Use structured processing technology to generate military element description data; By using a knowledge reasoning algorithm, relevant information is extracted from the military knowledge base based on the military element description data for comparison, thereby generating a comparison result; Inferring the effectiveness attributes of the military elements based on the comparison results; In the knowledge graph, the effectiveness attributes of the military elements are added to generate the military element knowledge graph.
8. A military document information extraction and knowledge graph construction system, characterized in that: include: A candidate text region generation module is used to obtain a military document PDF file, divide the text region according to the text layout characteristics, and generate a candidate text region set; A real text area screening module is used to match the candidate text area set with a military terminology knowledge base using a term matching algorithm to screen out real text areas; A module for generating candidate symbol annotation regions is used to perform shape analysis on non-text regions according to symbol shape features to generate a set of candidate symbol annotation regions; A real symbol annotation area screening module is used to match the candidate symbol annotation area set with a military symbol knowledge base using a pattern recognition algorithm to screen out real symbol annotation areas; A mapping relationship building module is used to determine the correlation between the real text area and the real symbol annotation area according to their relative positions, and to build a mapping relationship between the text and the symbol annotation; A text area correction module, for identifying text content using an OCR model for the text area for which a mapping relationship is established, and correcting the recognition result according to the military terminology knowledge base; A symbol annotation region correction module is used to obtain symbol annotation information using a symbol recognition model for the symbol annotation region for which a mapping relationship is established, and to correct the recognition result according to the military symbol knowledge base; A knowledge graph module is constructed to associate the corrected text recognition results and symbol annotation information according to the mapping relationship between the text and the symbol annotation, generate structured military element description data, and construct a military element knowledge graph.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the military document information extraction and knowledge graph construction method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: It includes computer programs and instructions. When the computer program or the instructions are run on a computer, the computer executes the military document information extraction and knowledge graph construction method as described in any one of claims 1 to 7.