YOLOv8-based bill document rotation target information detection and positioning method
Patent Information
- Application Number
- CN202610944154.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-18
AI Technical Summary
传统方案多依赖霍夫变换、投影校正等固定规则算法完成倾斜矫正,仅能适配小角度、高规整度的标准票据,面对任意角度倾斜、复杂表格纹理、轻微形变的票据图像极易出现校正失效、行列错乱、文本偏移等问题,直接导致OCR识别乱行、漏识、错识
[0009] Analysis of the YOLOv8-based method for detecting and locating rotating targets in invoice documents provided by this invention reveals that, in practical applications, this solution effectively addresses the technical challenges of poor tilt adaptability, single feature extraction, missing layout logic, and low character segmentation accuracy in traditional invoice detection through an integrated architecture of image standardization preprocessing, multi-scale rotation detection, layout structure parsing, and adaptive character segmentation. This significantly improves the robustness and accuracy of locating and segmenting complex invoice text. The solution normalizes the image size of the invoice by scaling proportionally and filling edges, avoiding image stretching distortion and content loss. Simultaneously, it normalizes pixel values and performs HWC and CHW dimension conversion, generating a standardized input tensor adapted to the deep learning framework. This step unifies the input specifications of invoice images from different shooting angles and resolutions, eliminating inference disturbances caused by inconsistencies in original image size, brightness, and dimensional format, and providing a standardized data foundation for subsequent stable feature extraction by the network.
Smart Images

Figure CN122780967A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of documents, and in particular to a method for detecting and locating rotational target information in YOLOv8 invoice documents. Background Technology
[0002] Automatic document recognition is a core technology for intelligent office scenarios such as financial archiving, document verification, intelligent reimbursement, and financial risk control. In actual data collection, document images are often captured by mobile phones and scanned from multiple angles, commonly exhibiting problems such as overall tilt, local offset, and irregular layout. Furthermore, documents contain complex layouts with numerous table rows and columns, interwoven fields, and mixed large and small characters, placing extremely high demands on the accuracy and robustness of text localization and character segmentation. Accurate, complete, and structured document character detection and segmentation are the prerequisites for ensuring the accuracy of subsequent OCR text recognition, field extraction, and content verification.
[0003] Currently, traditional invoice detection and recognition technologies have significant technical shortcomings. Traditional solutions mostly rely on fixed-rule algorithms such as Hough transform and projection correction to complete tilt correction, which can only adapt to standard invoices with small angles and high regularity. When faced with invoice images with arbitrary tilt angles, complex table textures, or slight deformations, they are prone to problems such as correction failure, row and column disorder, and text offset, directly leading to OCR recognition issues such as disordered lines, missed recognitions, and misidentifications. At the same time, most existing mainstream detection models adopt horizontal anchor box detection mechanisms, which can only fit horizontally and vertically arranged text regions. They cannot adapt to the geometric shapes of tilted text, diagonal cells, and offset fields in invoices, and are prone to defects such as inaccurate text box fitting, partial truncation, and unwanted background mixing.
[0004] Furthermore, existing deep learning detection methods have limited feature fusion capabilities, mostly relying on unidirectional top-down feature transfer. Shallow details and deep semantics cannot complement each other bidirectionally, making it difficult to simultaneously consider both minute character details and global layout features on the document. This results in significant issues of missed detections of small characters and false positives for large layouts. Conventional detection methods only output the coordinates of independent text boxes, lacking the ability to model the topological structure of the document layout. They cannot represent spatial relationships such as row and column arrangement, field adjacency, and table nesting, making it difficult to distinguish between detection fragments, adjacent fields, and cross-column intervals. This leads to problems such as multiple fields adhering to the same row, misaligned fields, and confused cell boundaries. Ultimately, during character segmentation, issues such as multiple character adhering, incomplete segmentation, and inconsistent sizes easily occur, severely reducing the overall accuracy and stability of automated document recognition. Summary of the Invention
[0005] The purpose of this invention is to provide a method for detecting and locating rotating target information in YOLOv8 document tickets, which solves the above-mentioned technical problems pointed out in the prior art.
[0006] This invention provides a method for detecting and locating rotational target information in invoice documents based on YOLOv8. The specific operation method includes: acquiring an image of the invoice document to be detected, normalizing the size of the image of the invoice document to be detected, and obtaining a standardized input tensor.
[0007] The standardized input tensor is input into the backbone network of YOLOv8, and a multi-layer feature map set is output. The multi-layer feature map set is analyzed according to spatial resolution to construct an enhanced feature map pyramid. Multiple sizes of rotational anchor boxes are set for the enhanced feature map pyramid and input into the detector head of YOLOv8 to output the target rotational bounding box. The target rotational bounding box is analyzed to obtain a single character image.
[0008] Compared with the prior art, the embodiments of the present invention have at least the following technical advantages:
[0009] Analysis of the YOLOv8-based method for detecting and locating rotating targets in invoice documents provided by this invention reveals that, in practical applications, this solution effectively addresses the technical challenges of poor tilt adaptability, single feature extraction, missing layout logic, and low character segmentation accuracy in traditional invoice detection through an integrated architecture of image standardization preprocessing, multi-scale rotation detection, layout structure parsing, and adaptive character segmentation. This significantly improves the robustness and accuracy of locating and segmenting complex invoice text. The solution normalizes the image size of the invoice by scaling proportionally and filling edges, avoiding image stretching distortion and content loss. Simultaneously, it normalizes pixel values and performs HWC and CHW dimension conversion, generating a standardized input tensor adapted to the deep learning framework. This step unifies the input specifications of invoice images from different shooting angles and resolutions, eliminating inference disturbances caused by inconsistencies in original image size, brightness, and dimensional format, and providing a standardized data foundation for subsequent stable feature extraction by the network.
[0010] Furthermore, this solution utilizes the YOLOv8 backbone network to extract multi-scale feature maps, combines spatial pyramid pooling to expand the receptive field of deep features, and constructs an enhanced feature pyramid through bidirectional upsampling and fusion of shallow and deep features. This ensures that each layer of features simultaneously possesses both shallow character detail information and global layout semantic information, accommodating the recognition needs of both tiny characters and large-sized table structures. This invention abandons the traditional fixed horizontal anchor frame mechanism, configuring multi-size, multi-aspect-ratio, and multi-angle rotating anchor frames in the feature pyramid to accurately adapt to unconventional text layouts such as slanted text on invoices and diagonal cells. Regression using the detector head yields a result that closely matches the true form of the text. Rotating bounding boxes solves the problems of inaccurate horizontal bounding box alignment, edge truncation, and background redundancy in traditional methods. Simultaneously, based on the target rotated bounding box, a document layout topology is constructed. Row and column field units are aggregated through row grouping, horizontal overlap analysis, and two-level spacing thresholds. Combined with horizontal and vertical projection curves, segmentation points are accurately extracted. Foreground density filtering and connected component denoising algorithms remove noise, artifacts, and interference. The final output is a single character image with uniform size, complete boundaries, and no redundant interference. This achieves fully automated integrated processing of tilted text localization, layout modeling, and character segmentation for documents, significantly improving the accuracy of document detection and localization in complex scenarios. Attached Figure Description
[0011] Figure 1 This is a flowchart of the main process of a YOLOv8-based method for detecting and locating rotating target information in a receipt document, as described in Example 1.
[0012] Figure 2 This is a flowchart of an enhanced feature map pyramid for a YOLOv8-based method for detecting and locating rotating target information in a ticket document, as described in Example 1.
[0013] Figure 3 This is a pooling diagram illustrating a YOLOv8-based method for detecting and locating rotating target information in a ticket document, as described in Embodiment 1.
[0014] Figure 4 This is a flowchart of the layout topology diagram of a YOLOv8-based method for detecting and locating rotating target information in invoice documents, as described in Embodiment 1.
[0015] Figure 5 The flowchart below shows a character image of a method for detecting and locating rotated target information in a YOLOv8 invoice document, as described in Embodiment 1.
[0016] Figure 6 This is a schematic diagram of the main flowchart and internal links of a YOLOv8-based method for detecting and locating rotating target information in a ticket document, as described in Embodiment 1.
[0017] Figure 7 This is a flowchart of a YOLOv8-based ticket document rotation target information detection and positioning system according to Embodiment 2;
[0018] Labels: Acquisition module 10; Analysis module 20. Detailed Implementation
[0019] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings.
[0021] Example 1
[0022] like Figure 1 As shown, this invention provides a method for detecting and locating rotational target information in YOLOv8 document receipts. The specific operation method includes:
[0023] S10: Obtain the image of the invoice document to be detected, and normalize the size of the image of the invoice document to be detected to obtain a standardized input tensor;
[0024] It should be noted that the process involves reading the image of the document to be inspected (which may be a color or grayscale image), obtaining the original width and height of the image, determining the ratio between the original size and the preset input size (e.g., 640×640 pixels), scaling the image proportionally to the preset size on the long side and padding the short side (usually with black pixels) to ensure that there is no compression or distortion or loss of content. The pixel values (0-255) of the scaled and padded image are converted to floating-point numbers and normalized to the 0-1 range. The data dimension order is adjusted from HWC (height, width, channel) format to CHW (channel, height, width) format. The processed data is then encapsulated into a tensor object under the PyTorch or TensorFlow framework to form a standardized input tensor.
[0025] S20: Input the normalized input tensor into the backbone network of YOLOv8 to output a multi-layer feature map set; construct an enhanced feature map pyramid based on the spatial resolution analysis of the multi-layer feature map set; set multiple sizes of rotational anchor boxes for the enhanced feature map pyramid and input them into the detector head of YOLOv8 to output target rotational bounding boxes; analyze the target rotational bounding boxes to obtain a single character image;
[0026] It should be noted that in the above steps, the standardized tensor is fed into the YOLOv8 backbone network to extract multi-scale original feature maps. Spatial pyramid pooling is used to expand the receptive field of deep features. Then, through bidirectional upsampling and downsampling of shallow and deep features, an enhanced feature map pyramid is constructed (i.e., the enhanced feature map pyramid specifically refers to the collection structure of shallow, middle, and deep feature maps generated after bidirectional cross-scale upsampling, downsampling, convolution, and transformation of multi-layer feature map sets; this structure differs from the native unidirectional feature pyramid of YOLOv8). Each layer of features simultaneously integrates shallow stroke details and global layout semantics, taking into account the recognition needs of small characters and the entire table layout of the document. Multiple sizes, aspect ratios, and angles of rotating anchor frames are configured at each level of the pyramid to adapt to tilted text and diagonal table cells on the document, and the detection head completes this process simultaneously. The process involves rotating bounding boxes for regression and classifying document fields, outputting target rotating bounding boxes that accurately reflect the actual tilt of the text and images. Subsequently, based on the coordinates of the rotating bounding boxes and local features, the document layout topology is constructed. Row and column fields are then divided using row grouping, horizontal overlap, and two-level spacing thresholds. Segmentation points are extracted using horizontal and vertical projection curves, and after foreground density filtering and connected component denoising, clean and well-organized single-character images are output. This entire processing solution achieves precise localization of tilted text on documents, spatial logic modeling of the layout, cell field aggregation, and character segmentation, effectively addressing the shortcomings of traditional horizontal detection boxes that cannot adapt to tilted documents, difficulties in separating overlapping characters, and loss of layout structure. It outputs clean, uniformly sized, and independent single-character images, providing high-quality pre-processing material for document OCR text recognition.
[0027] like Figure 2 As shown, specifically, in step S20, the normalized input tensor is input into the backbone network of YOLOv8, outputting a multi-layer feature map set; the multi-layer feature map set is analyzed according to spatial resolution to construct an enhanced feature map pyramid; multiple sizes of rotational anchor boxes are set for the enhanced feature map pyramid and input into the detector head of YOLOv8 to output the target rotational bounding box; the target rotational bounding box is analyzed to obtain a single character image. The specific operation steps are as follows:
[0028] S21: Input the normalized input tensor into the YOLOv8 backbone network (i.e., the backbone network contains multiple sequentially stacked convolutional blocks and multiple downsampling layers), and propagate the normalized input tensor forward through the multiple sequentially stacked convolutional blocks and multiple downsampling layers in the backbone network.
[0029] The normalized input tensor is downsampled by multiple downsampling layers, and the intermediate feature map corresponding to the downsampling layer is output. The spatial resolution (height × width) and number of channels (depth, which also represents the channel dimension) of each intermediate feature map are recorded.
[0030] Obtain the intermediate feature map output by the last downsampling layer in the backbone network, and use it as the highest layer feature map;
[0031] Spatial pyramid pooling is performed on the highest-layer feature map. (In traditional convolutional neural networks, fully connected layers require the input feature map to be of a fixed size (e.g., 7×7). This means that the network can only accept images of a fixed size (e.g., 640×640) during training. If the size of the input image changes, it needs to be forcibly cropped or stretched, which will lead to information loss or distortion. The purpose of spatial pyramid pooling is to allow the network to output a feature vector of a fixed length in the last layer, regardless of the size of the input image.) Multiple pooling windows with increasing sizes within the spatial pyramid are used to perform parallel pooling operations on the highest-layer feature map to obtain multiple pooling results.
[0032] The multiple pooling results are concatenated according to the number of channels to obtain a multi-scale pooling feature map;
[0033] It should be noted that the normalized input tensor is fed into the YOLOv8 backbone network (usually a CSPDarknet structure). The normalized input tensor propagates forward step by step, passing through convolutional blocks (containing convolution, batch normalization, and activation functions) and downsampling layers in sequence. Each downsampling layer halves the spatial size of the feature map (e.g., from 640×640 to 320×320) through convolution or pooling operations with a stride greater than 1, while gradually increasing the number of channels (e.g., from 64 to 128, 256, 512, etc.). After each downsampling layer completes downsampling, the intermediate feature map output of that layer is immediately extracted, and the spatial resolution (height × width) and number of channels (depth) of the intermediate feature map are recorded. These intermediate feature maps are arranged in the order of the downsampling layers and then fused for subsequent use.
[0034] Multi-scale pooling (spatial pyramid pooling) is performed on the feature map of the deepest layer (i.e., the bottommost and last layer) to enhance the receptive field. The highest-level feature map (with the smallest spatial resolution and the most channels) output from the last downsampling layer in the backbone network is located. Multiple pooling windows of increasing size (e.g., 5×5, 9×9, 13×13) are prepared, and pooling is performed on this highest-level feature map using pooling windows of different sizes simultaneously. Each pooling window operates independently, yielding pooling results with different receptive fields. The multiple pooling results are concatenated along the channel direction (not added) to form a new multi-scale feature map. This multi-scale feature map incorporates contextual information from pooling windows of different scales, such as... Figure 3 As shown;
[0035] In the above steps, the YOLOv8 native CSPDarknet backbone network is used to complete multi-scale basic feature extraction. The standardized input tensor is compressed layer by layer through multiple convolutional blocks and stride downsampling to expand the spatial size and feature channels. The intermediate feature maps output by each downsampling layer are retained simultaneously and the resolution and channel dimension are fully recorded. The three types of basic information, namely shallow fine texture, mid-level transition features and deep global semantics, are fully preserved. For the top and highest layer feature maps with the smallest resolution and weakest receptive field, this embodiment introduces multi-scale spatial pyramid parallel pooling. Multiple sets of pooling windows with increasing size are used to extract context information of different ranges simultaneously. The windows are spliced along the channel dimension to generate multi-scale pooled feature maps, which greatly expands the receptive field of deep features, makes up for the deficiency of insufficient context capture by a single fixed pooling window, strengthens the representation ability of long-distance related features such as document text, lines, and borders, and provides a global semantic basis for subsequent cross-level feature fusion.
[0036] S22: Sequentially obtain the intermediate feature maps output by each downsampling layer in the backbone network except for the last downsampling layer;
[0037] The multi-scale pooling feature map is upsampled to the same spatial resolution as each intermediate feature map, and the upsampled multi-scale pooling feature map is fused with each intermediate feature map element by element to obtain multiple fused feature maps.
[0038] Perform a convolution transformation on each fused feature map to obtain the hierarchical feature map corresponding to that downsampling layer, and combine the hierarchical feature maps corresponding to all downsampling layers with the highest layer feature map to form a multi-layer feature map set;
[0039] It should be noted that each intermediate feature map in the backbone network, except for the last downsampling layer, is extracted sequentially (processed one by one from deep to shallow). The multi-scale pooling feature map is upsampled (usually using nearest neighbor interpolation or bilinear interpolation). The target size of the upsampled feature map is kept consistent with the spatial resolution of the currently processed intermediate feature map. The upsampled multi-scale pooling feature map is then added element-wise with the current intermediate feature map (the pixel values at corresponding positions are added), which preserves high-level semantic information and supplements the spatial details of the shallow layers. A convolution operation (usually 1×1 or 3×3 convolution) is performed on the fused feature map to reduce the aliasing effect caused by fusion. The result of the convolution transformation is used as the feature map corresponding to that level and stored in the multi-layer feature map set. The above steps are repeated until all intermediate layer feature maps have been fused with high-level information. Finally, the highest layer feature map is also included in the multi-layer feature map set (as the top layer feature).
[0040] The above steps involve a two-way fusion of high-level semantics and shallow details, progressing from deep to shallow layers. Multi-scale pooled feature maps are upsampled level by level to match the resolution of each shallow intermediate feature map. Through element-wise addition, global layout semantics are injected while preserving shallow strokes and small character edge details, overcoming the limitation of shallow features only capturing local textures and lacking overall contextual constraints. After fusion, convolutional transformations eliminate feature aliasing interference caused by cross-scale splicing, regulate channel dimensions and feature distribution, and uniformly generate standardized hierarchical feature maps, which are then aggregated into a multi-layer feature map set. This constructs an original multi-scale feature system with complete deep and shallow layer information, solving the problem that single-scale features cannot simultaneously capture small character details and the overall layout of the ticket.
[0041] S23: Select the highest-level feature map in the multi-layer feature map set (find the feature map with the smallest spatial resolution (i.e., the deepest feature map) from the set, which is the highest-level feature map and the last downsampling layer) as the first reference feature map;
[0042] The fusion feature map with the second smallest spatial resolution in the multi-layer feature map set is selected as the first mid-scale layer feature map; the fusion feature map with the largest spatial resolution is selected as the bottom-scale layer feature map (shallow layer, richest in detail; where the second smallest spatial resolution refers to the fusion feature map with the second smallest spatial resolution in the multi-layer feature map set, and the smallest spatial resolution is the highest level feature map).
[0043] The first reference feature map is subjected to a first upsampling operation (i.e., enlarged size) according to the second smallest spatial resolution of the first mesoscale layer feature map to obtain the upsampled first reference feature map (i.e., the first reference feature map (top layer) is upsampled (enlarged size), and the target size of the upsampling is set to the spatial resolution of the first mesoscale layer feature map (middle layer); the upsampling method usually adopts bilinear interpolation or nearest neighbor interpolation. After the upsampling is completed, the spatial size of the first reference feature map is completely consistent with that of the first mesoscale layer feature map).
[0044] The upsampled first reference feature map and the first mesoscale layer feature map are concatenated along the number of channels (that is, concatenated along the channel direction) to obtain the first concatenated feature map (the number of channels of the concatenated feature map is equal to the sum of the number of channels of the two).
[0045] It should be noted that in the above steps, three core feature layers are divided from the multi-layer feature map set: top-level baseline, middle-level medium-scale, and bottom-level high-resolution, to build a top-down feature fusion link. The top-level baseline features are upsampled to the medium-level resolution and then stitched along the channels. Deep global semantics and medium-level table and field medium-scale structural features are integrated simultaneously. Through channel stitching, the two types of feature information are superimposed without loss, and the independent feature expressions of targets at different scales are completely preserved. This builds the basic fusion unit for the hierarchical progressive bidirectional enhancement pyramid and avoids the problem of feature information canceling out due to additive fusion.
[0046] S24: Perform a first convolution transformation on the first spliced feature map and a first intermediate fused feature map (perform a convolution operation on the spliced feature map (usually a 1×1 or 3×3 convolution). The purpose of convolution is to compress the number of channels, fuse the information of the two feature maps, and reduce aliasing).
[0047] A second upsampling operation is performed on the first intermediate fused feature map. The spatial resolution of the first intermediate fused feature map after the second upsampling is adjusted according to the maximum spatial resolution of the bottom-scale layer feature map (i.e., the target size of the second upsampling is set to the spatial resolution of the bottom layer feature map (shallow layer, maximum resolution). After the upsampling is completed, the spatial size of the adjusted first intermediate fused feature map is consistent with that of the bottom layer feature map). The first intermediate fused feature map after the second upsampling is then stitched together with the bottom layer feature map along the number of channels (i.e., along the channel direction) to obtain the second stitched feature map.
[0048] A second convolution transformation is performed on the second spliced feature map to obtain a second intermediate fused feature map, which is used as the first enhanced output feature map.
[0049] The first enhanced output feature map is downsampled for the first time. The spatial resolution of the first enhanced output feature map after the first downsampling is adjusted according to the spatial resolution of the first intermediate fusion feature map (that is, the second intermediate fusion feature map (shallow enhanced feature) is downsampled (reduced in size). The target size of the downsampling is set to the spatial resolution of the first intermediate fusion feature map (middle layer). The downsampling method usually adopts convolution or max pooling with a stride of 2. After the downsampling is completed, the spatial size of the second intermediate fusion feature map is consistent with that of the first intermediate fusion feature map). The first enhanced output feature map after the first downsampling is concatenated with the first intermediate fusion feature map along the number of channels (that is, concatenated along the channel direction) to obtain the third concatenated feature map.
[0050] The third convolution transformation is performed on the third spliced feature map to obtain the third intermediate fused feature map (middle layer re-enhanced feature), which is used as the second enhanced output feature map;
[0051] A second downsampling operation is performed on the second enhanced output feature map. The spatial resolution of the second enhanced output feature map after the second downsampling is adjusted according to the spatial resolution of the first reference feature map. (A second downsampling operation is performed on the third intermediate fusion feature map (middle layer re-enhancement). The target size of the downsampling is set to the spatial resolution of the first reference feature map (top layer / deep layer). After the downsampling is completed, the spatial size of the third intermediate fusion feature map is consistent with that of the first reference feature map.) The second enhanced output feature map after the second downsampling is then stitched together with the first reference feature map along the number of channels (that is, along the channel direction) to obtain the fourth stitched feature map.
[0052] The fourth convolution transformation is performed on the fourth spliced feature map to obtain the fourth intermediate fused feature map (deep re-enhanced feature), which is used as the third enhanced output feature map;
[0053] The first enhanced output feature map, the second enhanced output feature map, and the third enhanced output feature map are stacked in ascending order of spatial resolution to form a three-layer enhanced feature map pyramid, which serves as the enhanced feature map (that is, the first enhanced output feature map (shallow layer / highest resolution), the second enhanced output feature map (medium layer / medium resolution), and the third enhanced output feature map (deep layer / lowest resolution) are arranged in descending order of spatial resolution, in the order of shallow layer (highest), medium layer (medium), and deep layer (lowest). The three sorted feature maps are organized into a hierarchical structure (similar to a pyramid shape) as the enhanced feature map pyramid).
[0054] It should be noted that the above steps involve constructing a bidirectional feature enhancement loop that first samples from top to bottom and then from bottom to top. First, the middle-layer spliced features are upsampled a second time and spliced with the bottom-layer high-resolution features to generate a shallow enhanced output feature map, strengthening subtle target features such as extremely small single characters and fine dividing lines. Then, the shallow fine features are downsampled layer by layer and fed back to the middle and top layers to generate middle and deep enhanced feature maps respectively. Throughout the process, redundant channels are compressed and spliced feature noise is smoothed through multiple convolutional transformations. Finally, the shallow, middle, and deep enhanced feature maps are combined to construct an enhanced feature map pyramid. Compared to the native YOLOv8 single-layer unidirectional feature fusion, this step achieves bidirectional flow and mutual complementarity between shallow and deep features. Each layer of features simultaneously possesses fine-grained texture and global layout semantics, significantly improving the feature discrimination of characters and table borders of different sizes of tickets, and reducing the probability of missed detection of small characters and false detection of large layout boundaries.
[0055] S25: Each layer of the enhanced output feature map in the enhanced feature map pyramid is uniformly divided into grid cells. Multiple sizes of rotating anchor boxes are set according to the anchor point positions of the grid cells. These rotating anchor boxes are input into the regression branch of the YOLOv8 detection head to obtain predicted rotating bounding boxes. Furthermore, each rotating anchor box is analyzed simultaneously using the classification branch containing preset categories within the detection head to obtain target rotating bounding boxes. Adjacent target rotating bounding boxes are rotated, and the Euclidean distance between their center points is calculated. Adjacent bounding box pairs are filtered using the Euclidean distance between their center points. The connection relationship type of the adjacent bounding boxes is determined, and a layout structure topology graph of the document is constructed. Connected component analysis is performed on the layout structure topology graph to obtain individual character images.
[0056] It should be noted that in the above steps, each layer of the enhanced feature pyramid is uniformly divided into grids and configured with rotating anchor frames of various specifications to adapt to non-horizontal and non-vertical elements such as tilted text and diagonal dividing lines on the ticket. The input detection head simultaneously completes the rotation bounding box regression and character classification, accurately outputting the target rotation bounding box that fits the actual tilt angle of the character. The above steps filter adjacent character frames by Euclidean distance of the bounding box center points, and determine the connection relationship based on spatial distance and arrangement order to construct a layout topology map, completely restoring the spatial logical structure of the ticket fields, tables, rows and columns. Finally, the connected component algorithm is used to split the topology map into regions, accurately cutting out independent and unconnected single character images. This step achieves integrated processing of accurate positioning of tilted characters on the ticket, modeling of spatial relationships on the page, and automatic character segmentation, effectively solving the technical defects of traditional horizontal anchor frames that cannot adapt to tilted text on the ticket, difficulty in splitting connected characters, and loss of layout logic, providing clean, independent, and accurately positioned single character image materials for ticket text recognition.
[0057] like Figure 4 As shown, specifically, in step S25, each layer of the enhanced output feature map in the enhanced feature map pyramid is uniformly divided into grid cells, and multiple sizes of rotating anchor boxes are set according to the anchor point positions of the grid cells; the multiple sizes of rotating anchor boxes are input into the regression branch of the detection head in the YOLOv8 to obtain the predicted rotating bounding boxes; further, the classification branch containing preset categories in the detection head is used to analyze each rotating anchor box simultaneously to obtain the target rotating bounding box; adjacent target rotating bounding boxes are rotated and the Euclidean distance between their center points is calculated; the Euclidean distance between their center points is used to filter adjacent bounding box pairs; the connection relationship type of the adjacent bounding boxes is determined, and the layout structure topology of the document is constructed; the layout structure topology is analyzed using a connected component algorithm to obtain a single character image. The specific operation steps are as follows:
[0058] S251: For each layer of enhanced output feature map in the enhanced feature map pyramid, the enhanced output feature map of that layer is evenly divided into grid cells, and each grid cell corresponds to an anchor point position (i.e., the center point of the grid cell).
[0059] Multiple rotating anchor frames of different specifications are defined using the x-coordinate, y-coordinate, width, and height parameters of the center point (the center point position is the anchor point location) and a preset initial rotation angle parameter at each anchor point location. Each rotating anchor frame has different size parameters (the x-coordinate and y-coordinate positions, reflecting large, medium, and small dimensions), aspect ratio parameters (width and height parameters, indicating a slender, square, or wide and flat shape), and rotation angle parameters (such as 0°, 30°, and 60°). Each rotating anchor frame's specifications are also defined by the x-coordinate, y-coordinate, width, height, and rotation angle parameters of its center point (anchor point location). Multiple anchor frames at the same anchor point location share the same center point coordinates, but their width, height, and angle differ. The initial rotation angle parameter is the initial angle set by the anchor frame itself, while the rotation angle parameter is the actual value used to describe the orientation of the frame. Furthermore, all rotating anchor frames at the same anchor point location share the feature vector at that anchor point location.
[0060] Based on each rotated anchor box at the anchor point position in the enhanced output feature map of each layer, extract the feature vector of the anchor point position corresponding to each rotated anchor box (that is, all channel dimensions in the grid cell at the anchor point position, i.e., the number of channels. The channel dimensions can be obtained through image stitching and convolution in the above steps, so they will not be repeated).
[0061] The feature vector of the anchor point position corresponding to each rotating anchor frame is input into the regression branch of the YOLOv8 detector head. The regression branch outputs the horizontal coordinate offset of the center point, the vertical coordinate offset of the center point, the width scaling, the height scaling (i.e., the width scaling and height scaling represent the changes in the size of the rotating anchor frame, such as a reduction of 6 pixels), and the rotation angle offset of the rotating anchor frame. Although all rotating anchor frames at the same anchor point position share the feature vector at that anchor point position, the feature vector is a "one-to-many" input. This feature vector will be copied and used for each rotating anchor frame at that position, so that the regression branch and classification branch of the detector head can perform parallel calculations. This allows for the analysis of all rotating anchor frames at that anchor point position, thereby enabling the regression branch of the YOLOv8 detector head to more accurately output all offsets and the type of analysis in the classification branch.
[0062] It should be noted that in the above steps, the feature map of each layer of the enhanced feature pyramid is uniformly divided into grid units, and multiple sizes of rotating anchor frames are matched with the grid center as the anchor point. The anchor frames are configured with different center points, width and height dimensions, aspect ratios and initial rotation angles, which can fully adapt to various types of text and table cells in the document, including horizontal, tilted and vertical ones. Each anchor point reuses the complete channel feature vector of the corresponding grid and is simultaneously fed into the regression branch of the detection head. The network learns the offset correction amount of the center, width and height and angle of various rotating anchor frames, breaking the limitation that the traditional horizontal anchor frame cannot adapt to tilted document text. It covers the full angle and full size text targets of the document in advance, and provides sufficient prior anchor base for accurate regression of rotating frames.
[0063] S252: Sum the horizontal coordinate offset of the center point with the horizontal coordinate of the center point of the rotating anchor frame to obtain the predicted horizontal coordinate of the center point;
[0064] The sum of the offset of the center point's ordinate and the ordinate of the center point of the rotating anchor frame is used as the predicted center point's ordinate.
[0065] The predicted width is obtained by multiplying the width scaling amount by the width parameter of the rotating anchor frame.
[0066] The predicted height is obtained by multiplying the height scaling amount by the height parameter of the rotating anchor frame.
[0067] The rotation angle offset is summed with the rotation angle parameter of the rotating anchor frame to obtain the predicted rotation angle.
[0068] A predicted rotation bounding box is constructed based on the predicted center point x-coordinate, predicted center point y-coordinate, predicted width, predicted height, and predicted rotation angle (i.e., using the center point x-coordinate of each anchor point position). ), center point ordinate ( ), preset width ( ), preset height ( ) and preset rotation angle ( This is used to define multiple specifications of rotating anchor frames; among them, The range of values is This is used to set the initial orientation of the anchor frame; the feature vector corresponding to each rotated anchor frame is input into the regression branch of the detector head in YOLOv8 to construct the predicted rotated bounding box, and the x-coordinate offset of the center point ( ), center point ordinate offset ( Width scaling () ), height scaling ( ) and rotation angle offset ( The parameters of the predicted rotated bounding box are calculated using the following formula, and the x-coordinate of the predicted center point is calculated. Predicted center point ordinate: Prediction width: (or The method chosen depends on the actual regression method; here, we take exponential transformation as an example, representing scale scaling (where Δw is also the width offset). Predicted height: (in Also expressed as height offset); predicted rotation angle: (That is, the sum of the preset angle of the anchor frame and its offset); based on the calculated... Five parameters are used to construct the final predicted rotated bounding box.
[0069] It should be noted that in the above steps, the anchor frame coordinates, size, and rotation angle are corrected based on the five types of offsets output by the regression branch. Through the calculation logic of coordinate addition, size scaling, and angle superposition, the preset coarse-grained rotation anchor frame is corrected to the predicted rotation bounding box that fits the real text area. The slanted geometric features of the document text are fully preserved, and the actual enclosing range of slanted text and diagonal table lines is accurately restored. This avoids the target box fitting deviation caused by fixed-angle anchor frames and ensures that the geometric contour positioning of each text area is accurate.
[0070] S253: Using the classification branch containing preset categories within the detection head, the feature vectors at the anchor points corresponding to each rotating anchor frame are analyzed simultaneously, outputting the confidence score vectors for each preset category corresponding to the rotating anchor frame (i.e., each channel in the feature vector corresponds to a preset category, such as invoice number, date, amount, name, and other invoice field types); the confidence score vector with the maximum value is selected as the predicted confidence score of the rotating anchor frame, and the preset category corresponding to the predicted confidence score is taken as the predicted category of the rotating anchor frame; and the predicted rotating bounding box corresponding to the predicted confidence score is extracted (i.e., the feature vectors are analyzed simultaneously using the regression and classification branches of the detection head to obtain the predicted rotating bounding box and the predicted category; and because of the simultaneous analysis, the predicted confidence score also corresponds to the predicted rotating bounding box constructed by the rotating anchor frame).
[0071] If the predicted confidence score is greater than or equal to the preset confidence threshold (i.e., 0.3 ~ 0.7, preferably 0.5), then the predicted rotated bounding box is taken as a candidate rotated bounding box (traverse all rotated anchor boxes at all anchor points of all enhanced output feature maps to obtain a set of candidate rotated bounding boxes).
[0072] Calculate the rotation intersection-union ratio (ROU) between any two candidate rotated bounding boxes. The ROU is calculated by dividing the area of the intersection of the two candidate rotated bounding boxes by the area of their union; that is, by calculating the area of the overlapping region (intersection) of the two rotated rectangles, subtracting the area of the intersection (union) from the sum of the areas of the two rectangles, and then dividing the intersection area by the union area to obtain the ROU value (range 0~1, a larger value indicates more severe overlap). The ROU is calculated by representing the two candidate rotated bounding boxes as follows: and and will and The coordinates of the four corner points are converted into planar point sets respectively. and The polygon clipping algorithm is used to clip the polygonal region of R1 using the four edges of R2, and the vertex set of the overlapping region is extracted. ;like If empty, then the rotation-intersection-union ratio is determined to be 0; if not empty, the area of the overlapping region is calculated using the shoelace formula. Fourth step, calculate the area of the union. - Step 5: Calculate the rotation-intersection-union ratio. ;
[0073] If the rotation crossover ratio (CCR) is greater than a preset CCR threshold (i.e., 0.4 ~ 0.6, preferably 0.45), then the candidate rotation bounding box with the highest predicted confidence score is selected from any two candidate rotation bounding boxes. The process is iterated through all selected candidate rotation bounding boxes until the CCR between any two selected candidate rotation bounding boxes is less than or equal to the CCR threshold. All the candidate rotation bounding boxes that are finally retained are then used as target rotation bounding boxes (i.e., the predicted confidence scores of two candidate rotation bounding boxes are compared, the candidate rotation bounding box with the lower predicted confidence score is removed, the candidate rotation bounding box with the higher predicted confidence score is retained, and the process of calculating the CCR between any two candidate rotation bounding boxes continues until the CCR is less than or equal to the CCR threshold. The iteration is then stopped, and the last remaining candidate rotation bounding box is used as the target rotation bounding box).
[0074] It should be noted that in the above steps, the detection head regression and classification branch are performed in parallel, and the predicted bounding boxes and the confidence scores of various document fields are output simultaneously. Only candidate boxes with confidence scores higher than the confidence threshold are retained, and low-confidence noise and false targets with messy textures are filtered out. The rotation intersection-union ratio non-maximum suppression algorithm is used to iteratively remove low-confidence overlapping boxes based on the overlapping area of the rotating rectangle, eliminating the problem of repeated detection of multiple boxes in the same text region, and outputting a set of target rotating bounding boxes without redundancy and overlap, which greatly reduces the redundant computation of subsequent page topology construction and purifies the effective text boxes.
[0075] S254: Determine the orientation of the target rotation bounding box in the image of the document to be detected based on the predicted rotation angle of the target rotation bounding box (i.e., the predicted rotation angle is represented by the degree of tilt; the tilt direction of the target rotation bounding box in the image plane is determined based on the value of the predicted rotation angle; for example: an angle close to 0° indicates that it is close to horizontal, an angle close to 90° indicates that it is close to vertical, and a negative angle indicates that it is tilted to the left).
[0076] The spatial location of the target rotated bounding box is determined based on the predicted center point x-coordinate and predicted center point y-coordinate (i.e., the specific position of the box is located in the image coordinate system using the center point x-coordinate (horizontal position) and y-coordinate (vertical position), and the coverage area of the target rotated bounding box, i.e., the spatial location, is determined by combining the width and height).
[0077] Using the spatial positioning of the target rotating bounding box as the rotation center, a rotation transformation (i.e., reverse rotation) is performed according to the negative angle value of the predicted rotation angle, so that the orientation of the target rotating bounding box after the rotation transformation is corrected to the horizontal direction, and a rectangular local region is cropped from the enhanced output feature map after the rotation transformation according to the predicted width and predicted height.
[0078] A pooling operation is performed on the cropped rectangular local region (i.e., adaptive pooling is performed on the rectangular local region, and the pooling layer transforms it into a preset fixed spatial size (e.g., 7×7 pixels) regardless of the input region size. This process does not change the number of channels, but only unifies the height and width). The rectangular local region is pooled to the preset fixed spatial size to obtain a local region feature map of the fixed spatial size (after pooling, a local region feature map of uniform size (e.g., 7×7×C) is obtained). The local region feature map is unfolded into a one-dimensional feature vector, which serves as the local region feature vector corresponding to the target rotated bounding box (the fixed-size feature map is fully unfolded in both spatial and channel dimensions and flattened into a one-dimensional vector, which is the local region feature vector corresponding to the rotated bounding box, containing the visual semantic information of the text region in the image to be detected).
[0079] Based on the predicted x-coordinate and y-coordinate of the center point of all target rotated bounding boxes, calculate the Euclidean distance between the center points of every two target rotated bounding boxes in a two-dimensional plane coordinate system (that is, the Euclidean distance is the straight-line distance between two points, reflecting the spatial distance between the two boxes).
[0080] It should be noted that in the above steps, the target box is corrected to a horizontal orientation based on the predicted rotation angle, the local image is cropped and the feature size is unified through adaptive pooling, and a standardized one-dimensional local feature vector is generated by straightening. This eliminates the differences in feature scale and angle caused by the tilt of the document text and unifies the feature expression of text regions of different sizes and tilt angles. The steps calculate the Euclidean distance between the pairwise target boxes through the center point coordinates to quantify the spatial relationship between text regions and provide a quantitative spatial index for determining adjacent text units and constructing page layout relationships.
[0081] S255: If the Euclidean distance between the center points is less than a preset distance threshold (i.e., image width × 0.05 ~ 0.15), then the two target rotation bounding boxes are combined to form an adjacent bounding box pair. All rotation bounding box combinations are traversed to obtain all adjacent bounding box pairs (i.e., two target rotation bounding boxes, one of which is the first target rotation bounding box and the other is the second target rotation bounding box, for easy differentiation).
[0082] Obtain the first local region feature vector corresponding to the first target rotated bounding box and the second local region feature vector corresponding to the second target rotated bounding box in the adjacent bounding box pair;
[0083] Calculate the difference between the feature vector of the first local region and the feature vector of the second local region, and use it as the lookup vector (i.e., reflecting the degree of difference between the features of the two regions).
[0084] The feature vectors of the first local region and the second local region are summed to obtain a sum vector (which reflects the commonalities and comprehensive information of the features of the two regions).
[0085] The difference vector and the sum vector are concatenated according to the channel dimension (number of channels) to obtain the concatenation relationship feature vector;
[0086] The concatenated relationship feature vector is input into a pre-trained relationship prediction network, which outputs the connection relationship type. This connection relationship type includes horizontal adjacency (two boxes in the same row, arranged horizontally (like different fields in the same row)), vertical adjacency (two boxes in the same column, arranged vertically (like the same field in different rows)), and containment (one box contains another box (like a table frame containing a cell)). The relationship prediction network contains multiple fully connected layers (i.e., first fully connected layer: input dimension D = 25088, output dimension 1024; second fully connected layer: 1024 -> 256; third fully connected layer: 256 -> ...). 3) The concatenation relationship feature vector is nonlinearly transformed by the multiple fully connected layers and nonlinear activation layers (i.e., the first and second fully connected layers are followed by ReLU activation functions; the third fully connected layer is followed by a batch normalization layer without ReLU; the output layer uses the Softmax activation function). The probability distribution of the adjacent bounding box pairs belonging to each preset connection relationship type is output (i.e., a three-dimensional probability distribution vector, with the three dimensions corresponding to the probability values of the three preset connection relationship types. This relationship prediction network is common knowledge and will not be elaborated further). The connection relationship type with the highest probability value is selected from the probability distribution as the connection relationship determination result of the adjacent bounding box pairs.
[0087] Each target rotated bounding box is used as a topology node, and the connection relationship type of all adjacent bounding box pairs is used as the edge attribute between the corresponding topology nodes to construct the layout structure topology graph of the ticket document (that is, each topology node in the layout structure topology graph contains the geometric position information (i.e., the Euclidean distance between the centers) and the local region feature vector of the target rotated bounding box corresponding to the topology node, and each edge attribute contains the connection relationship type of the relative position between the two topology nodes corresponding to the edge attribute).
[0088] It should be noted that in the above steps, adjacent bounding box pairs are filtered using a distance threshold. Two sets of text region features are fused to construct difference vectors and sum vectors, which are then concatenated into relational features. These features are fed into a pre-trained relation prediction network to automatically identify three types of layout relationships: horizontal adjacency, vertical adjacency, and inclusion. A layout topology graph is constructed using each rotated text box as a topology node and the relationship type between boxes as an edge attribute. This graph fully carries the global spatial logic of the ticket's row and column arrangement, table nesting, and field adjacency, achieving structured modeling of the ticket layout and distinguishing complex layout relationships such as ordinary text in the same row, corresponding fields above and below, and nested table cells.
[0089] S256: Based on the analysis of the predicted center point x-coordinate and predicted center point y-coordinate of the target rotated bounding boxes in the layout structure topology diagram, obtain row groups of topology nodes; determine the horizontal projection interval by obtaining the left and right boundary x-coordinate values of the target rotated bounding boxes in the row groups; construct a horizontal overlap relationship diagram for the horizontal projection intervals of any two target rotated bounding boxes in the row groups; perform connected component analysis on the horizontal overlap relationship diagram to obtain an ordered bounding box sequence within rows; extract two adjacent target rotated bounding boxes in the ordered bounding box sequence within rows to obtain field unit groups; construct a text unit image of the document from multiple field unit groups; perform binarization analysis on the text unit image to extract candidate vertical and candidate horizontal segmentation points; segment the text unit image using the candidate vertical and candidate horizontal segmentation points to obtain target character blocks; extract the circumscribed rectangle boundary of the foreground pixels of the target character blocks, and crop the circumscribed rectangle boundary of the corresponding position inside the target character block to obtain a standardized character image; perform connected component analysis on the standardized character image to obtain a single character image.
[0090] It should be noted that in the above steps, text lines are grouped based on the coordinates of nodes in the topology graph, and intra-line overlapping relationships are constructed based on the horizontal projection interval. An ordered text sequence is output through the connected component algorithm, and independent field unit groups are divided to generate text unit images. The next step extracts horizontal and vertical segmentation points through binarization to segment the text units into independent character blocks, extracts the bounding rectangle of the foreground pixels for standardized cropping, and finally outputs single character images with no adhesion and regular size through connected component segmentation. Based on the complete rotation correction and page topology constraints mentioned above, the problem of character segmentation disorder caused by slanted text adhesion, row and column mixing, and table nesting on the invoice is effectively solved. The segmented single character has complete boundaries and no extra background interference, and can be directly supplied to the invoice text recognition module.
[0091] like Figure 5As shown, specifically, in step S256, based on the analysis of the predicted center point x-coordinate and predicted center point y-coordinate of the target rotating bounding box in the layout structure topology diagram, row groups of topology nodes are obtained; the horizontal projection interval is determined by obtaining the left and right boundary x-coordinate values of the target rotating bounding box in the row group; a horizontal overlap relationship diagram is constructed for the horizontal projection intervals of any two target rotating bounding boxes in the row group; a connected component algorithm is performed on the horizontal overlap relationship diagram to obtain an intra-row ordered bounding box sequence; adjacent target rotating bounding boxes in the intra-row ordered bounding box sequence are extracted and analyzed to obtain field unit groups; a text unit image of the document is constructed from multiple field unit groups; the text unit image is binarized to extract candidate vertical segmentation points and candidate horizontal segmentation points; the text unit image is segmented using the candidate vertical segmentation points and candidate horizontal segmentation points to obtain target character blocks; the circumscribed rectangle boundary of the foreground pixels of the target character block is extracted, and the circumscribed rectangle boundary of the corresponding position is cropped inside the target character block to obtain a standardized character image; a connected component analysis is performed on the standardized character image to obtain a single character image. The specific operation steps are as follows:
[0092] S2561: Obtain the x-coordinate and y-coordinate of the predicted center point of the target rotated bounding box of all topological nodes in the layout structure topology diagram, and arrange all topological nodes in ascending order according to the value of the predicted center point y-coordinate to obtain the y-coordinate sorted node sequence (i.e., the node with the smallest y-coordinate value is ranked first (top of the image), and the node with the largest y-coordinate value is ranked last (bottom of the image)).
[0093] The first topological node in the sorted node sequence is selected as the starting node of the current row group, and the y-coordinate of the predicted center point of the starting node is marked as the reference y-coordinate of the current row.
[0094] Iterate through all the topological nodes (i.e., subsequent nodes) in the sorted node sequence by vertical coordinate, and calculate the difference between the vertical coordinate of the predicted center point of the currently traversed topological node and the vertical coordinate of the current row reference vertical coordinate.
[0095] If the difference in the ordinate is less than the preset row height threshold (i.e., 32 ~ 64), then the current topology node (subsequent node) is assigned to the current row group, and the current row reference ordinate is updated to the average of the predicted center point ordinates of all topology nodes in the current row group.
[0096] If the difference in the ordinate is greater than or equal to the row height threshold, the current row group is closed, and a new row group is created with the current topology node (subsequent node) as the starting node of the new row group. The ordinate of the predicted center point of the starting node of the new row group is marked as the reference ordinate of the new row. The topology nodes (i.e., subsequent nodes) are traversed until all topology nodes in the ordinate sorting node sequence have been traversed, resulting in multiple row groups.
[0097] Get the left and right x-coordinates of the target rotated bounding box for all topological nodes in the current row group;
[0098] The horizontal projection interval of each target's rotated bounding box is determined based on the left and right boundary x-coordinate values (i.e., the left boundary x-coordinate is used as the starting point of the interval, and the right boundary x-coordinate is used as the ending point of the interval, forming a horizontal projection interval [left boundary, right boundary]).
[0099] It should be noted that in the above steps, global top-down grouping is completed using the vertical coordinates of the topological nodes. Different text lines are dynamically distinguished based on the row height threshold. The grouping process continuously updates the row baseline vertical coordinate to the mean value within the group, adapting to scenarios where the height of text within the document is slightly offset or the vertical misalignment of tilted text. This avoids errors in splitting text within the same row or merging text across rows caused by a single fixed vertical coordinate threshold. At the same time, the left and right boundaries of all rotated text boxes within each row are extracted to generate standardized horizontal projection intervals, quantifying the horizontal coverage of each text unit. This provides a unified quantitative basis for calculating the left and right overlap and spacing of text within the same row, achieving fully automatic and accurate classification of document text lines and distinguishing different row fields.
[0100] S2562: Calculate the overlap length for the horizontal projection intervals of any two target rotated bounding boxes within a row group (i.e., for any two target rotated bounding boxes within a row group, extract their horizontal projection intervals and calculate the overlap length of the two intervals in the horizontal direction; the overlap length is calculated by taking the larger start point and the smaller end point of the two intervals. If the larger start point is smaller than the smaller end point, then the overlap length = smaller end point - larger start point; otherwise, the overlap length is 0); filter the minimum predicted width among the two target rotated bounding boxes (i.e., find the box with the smaller predicted width value from the two target rotated bounding boxes, divide the overlap length by the smaller predicted width to obtain the overlap ratio, which reflects the degree to which one box is covered by another box in the horizontal direction).
[0101] Calculate the overlap ratio between the overlap length and the predicted width of the minimum value (i.e., divide the overlap length by the predicted width of the minimum value to obtain the overlap ratio).
[0102] If the overlap ratio is greater than a preset overlap ratio threshold (i.e., 0.3 ~ 0.6, preferably 0.4), then it is determined that there is a horizontal overlap relationship between any two target rotated bounding boxes, and the two target rotated bounding boxes are recorded as a horizontal overlap pair; a horizontal overlap relationship graph within the row group is constructed based on all horizontal overlap pairs (i.e., each target rotated bounding box (topological node) within the row group is used as a vertex of the horizontal overlap relationship graph, and the line connecting each pair of boxes with a horizontal overlap relationship is used as an edge of the horizontal overlap relationship graph, thus forming a horizontal overlap relationship graph within the row group).
[0103] Perform connected component analysis on the horizontal overlap graph to merge the target rotated bounding boxes with horizontal overlap into the same overlap group, resulting in multiple overlap groups within that row group (i.e., a connected component is a set of vertices in the graph that are connected to each other (directly or indirectly through paths). All rotated bounding boxes belonging to the same connected component are merged into the same overlap group; the boxes within each overlap group have direct or indirect overlap in the horizontal direction, usually belonging to the same entity region, such as: a complete field being mistakenly cut into multiple detection boxes).
[0104] For each overlapping group, the target rotated bounding box is arranged in ascending order according to the left boundary x-coordinate value to obtain an ordered subsequence within the overlapping group;
[0105] For the target rotated bounding boxes within the row group that are not covered by the horizontal overlap pairs, they are taken as isolated rotated bounding boxes; and sorted in ascending order according to the left boundary x-coordinate values to obtain independent subsequences (that is, find the rotated bounding boxes within the row group that are not covered by any horizontal overlap pairs (i.e., isolated boxes that do not overlap with other boxes), extract these isolated boxes as independent subsequences, and sort them in ascending order according to the left boundary x-coordinate values).
[0106] The ordered subsequences within the overlapping group and the independent subsequences are merged and sorted according to their respective left boundary x-coordinate values to obtain the inline ordered bounding box sequence corresponding to the row group (that is, the ordered subsequences within the overlapping group and the independent subsequences are merged and sorted according to the overall distribution range of their respective left boundary x-coordinate values, i.e., arranged from left to right, to ensure that the position order of each box in the final sequence conforms to the left-right order in the actual page layout. After merging, the inline ordered bounding box sequence corresponding to the row group is obtained. Each position in the inline ordered bounding box sequence corresponds to a target rotated bounding box, arranged from left to right).
[0107] It should be noted that in the above steps, the overlap ratio criterion is constructed based on the ratio of the horizontal projection interval overlap length to the box width. This accurately identifies multiple fragmented text boxes with the same field generated by detection and segmentation. An undirected graph is constructed based on the overlap relationship, and the same source text fragments are merged through the connected component algorithm to restore the horizontal coverage of the complete field. The same source fragmented boxes with horizontal overlap are distinguished from independent text boxes with no intersection. Ordered subsequences are generated separately and then sorted by the horizontal coordinate. The output is an inline ordered boundary box sequence that strictly fits the layout of the ticket from left to right. This eliminates the deviation in left and right layout judgment caused by the tilt of the rotated box, unifies the horizontal spatial order of the text in the same row, and provides ordered input for field spacing and column division calculation.
[0108] S2563: Traverse the ordered bounding box sequence within each row group, extract the target rotation bounding box on the left of two adjacent target rotation bounding boxes as the left bounding box, and the target rotation bounding box on the right as the right bounding box (that is, take the first and second boxes from the sequence as the first pair of adjacent boxes, take the second and third boxes as the second pair, and so on, only two adjacent boxes are processed each time, and the box on the left of the two adjacent boxes is recorded as the left bounding box, and the box on the right is recorded as the right bounding box).
[0109] Extract the x-coordinate value of the right boundary of the left bounding box as the right edge position of the left bounding box; extract the x-coordinate value of the left boundary of the right bounding box as the left edge position of the right bounding box;
[0110] Calculate the absolute value of the difference between the right edge position of the left bounding box and the left edge position of the right bounding box, and use it as the horizontal distance between two adjacent target revolved bounding boxes (i.e., calculate the left edge position of the right bounding box minus the right edge position of the left bounding box, and take the absolute value of the difference (since the left box is to the left of the right box, the difference is usually positive, and taking the absolute value ensures safety). The absolute value of this difference is the horizontal distance between two adjacent target revolved bounding boxes).
[0111] The predicted width of the left bounding box is used as the width of the left bounding box, and the predicted width of the right bounding box is used as the width of the right bounding box; the ratio of the horizontal spacing to the width of the left bounding box is calculated as the left spacing ratio (that is, the horizontal spacing is divided by the predicted width of the left bounding box to obtain the ratio of the spacing to the width of the left box).
[0112] Calculate the ratio of the horizontal spacing to the width of the right bounding box, and use it as the right spacing ratio (i.e., divide the horizontal spacing by the predicted width of the right bounding box to obtain the ratio of the spacing to the width of the right box).
[0113] The ratio of the minimum value selected from the left spacing ratio and the right spacing ratio is used as the effective spacing ratio between two adjacent target rotated bounding boxes (that is, the reason for choosing a smaller ratio is that it is more conservative to use a narrower target rotated bounding box as a reference, so as to avoid the spacing ratio being underestimated due to a target rotated bounding box being too wide).
[0114] The preset first spacing threshold (0.1 ~ 0.3, the smaller value, used to distinguish between the same field and different fields) and the second spacing threshold (0.5 ~ 0.8, the larger value, used to distinguish between ordinary field spacing and cross-column spacing); that is, the second spacing threshold is greater than the first spacing threshold;
[0115] Determine whether the effective spacing ratio is less than the first spacing threshold (i.e., the first and second spacing thresholds are dimensionless ratios, and the effective spacing ratio of the comparison object (i.e., the ratio of horizontal spacing to the width of the target rotated bounding box) is also a dimensionless ratio; the use of dimensionless spacing ratios aims to eliminate the scale influence of different character widths and different image resolutions on the spacing determination results, and ensure the robustness and generalization of spacing determination).
[0116] If so, the two adjacent target rotation bounding boxes are marked as internal links; this means they belong to the same field unit within the same cell (e.g., "name" as a complete word in a table is the same field unit). The two adjacent target rotation bounding boxes are assigned the same field unit identifier, and subsequently merged into the same field unit group (this group must contain at least these two boxes). When two adjacent boxes are determined to belong to the same field unit (effective spacing ratio < first spacing threshold), the two boxes are assigned the same field unit identifier (e.g., both set to ID=5), indicating that the two adjacent target rotation bounding boxes belong to the same "field unit group" (i.e., the same cell / field content, which may have been segmented into multiple boxes due to detection). Figure 6 (as shown)
[0117] If not, then determine whether the effective spacing ratio is less than the second spacing threshold;
[0118] If so, then the two adjacent target rotated bounding boxes are marked as field intervals; this indicates that they belong to different independent field units within the same row (i.e., the field unit identifier assigned to the left bounding box is different from the field unit identifier assigned to the right bounding box; when it is determined that two adjacent boxes belong to different independent field units (effective spacing ratio ≥ first spacing threshold), the field unit identifier assigned to the left bounding box is set to be different from the field unit identifier assigned to the right bounding box (e.g., left box ID=5, right box ID=6), indicating that the two adjacent target rotated bounding boxes belong to different field unit groups; the field interval position refers to the boundary position between two adjacent rotated bounding boxes (detection boxes) within the same row that is determined to belong to different fields; among them, the field interval also represents the interval between adjacent field units within the same row (e.g., the interval between "name" and "gender"), which is the "row" interval).
[0119] If not, then the two adjacent target rotation bounding boxes are marked as column intervals; this means that there is a large column gap between the two adjacent boxes, they belong to different columns (and there is a cross-column gap, because the effective gap ratio is greater than or equal to the second gap threshold, so it means that the gap between the two adjacent target rotation bounding boxes is large), (that is, the column division priority of column intervals is higher than that of field intervals; column intervals are represented by adjacent field units in the same row and there is a cross-column gap (indicating that a column of content is missing in the middle, which belongs to a more obvious column boundary), that is, "column" intervals).
[0120] It should be noted that the above steps introduce a relative spacing ratio instead of a fixed pixel spacing threshold, using the width of the narrow frame in adjacent frames as a reference standard to avoid distortion in spacing determination caused by differences in text frame width; two levels of spacing thresholds are set to differentiate between continuous characters within cells, adjacent fields in the same row, and large-spaced fields across columns, and the two types of spacing positions are marked simultaneously to accurately distinguish between character gaps within cells, ordinary field gaps, and table column splitting gaps, and to quantitatively define the boundary positions of field units, thereby automatically dividing the same field within a cell into independent fields between rows, and accurately identifying the boundaries of invoice table columns and cells;
[0121] S2564: Merge two adjacent target rotated bounding boxes marked as internally connected within the ordered bounding box sequence to obtain an initial character fragment; use the field interval and column interval as mandatory delimiters to divide the ordered bounding box sequence into multiple consecutive fragment groups, which are multiple field unit groups (i.e., using the position marked as "field interval" or "column interval" as a mandatory delimiter, the entire in-line sequence, including the initial character fragment and unmarked isolated target rotated bounding boxes, is divided into multiple consecutive fragment groups); that is, the field unit groups contain adjacent field units within the same row, not the same group of fields within a single cell;
[0122] Get the left boundary x-coordinate value of all target rotated bounding boxes in each field cell group, filter the minimum left boundary x-coordinate value, and use it as the column starting boundary of that field cell group;
[0123] Get the right boundary x-coordinate value of all target rotated bounding boxes in each field cell group, filter the right boundary x-coordinate value of the maximum value, and use it as the column termination boundary of the field cell group;
[0124] The horizontal span interval of each field unit group is determined based on the column start boundary and the column end boundary (that is, it is represented as the total horizontal coverage of a field unit group, which is jointly determined by the column start boundary and the column end boundary).
[0125] Calculate the difference between the column start boundary and the column end boundary of any two field cell groups;
[0126] If the difference between the starting boundaries of the columns is less than a preset alignment threshold (10 to 20 pixels) and the difference between the ending boundaries of the columns is less than a preset alignment threshold, then any two field units are combined into the same column group to obtain the text unit image of the document (i.e., the text unit image includes the row index, column index, and corresponding target rotation bounding box set of each cell).
[0127] It should be noted that in the above steps, independent field unit groups are divided row by row based on field intervals and column intervals. The horizontal start and end boundaries of each group are extracted to generate the horizontal span interval of the field. The vertical fields in the same column are determined by the left and right boundary difference alignment threshold. The aggregation of multiple rows and columns of fields is completed to generate a complete text unit image. This step integrates the two-layer spatial logic of the bill row and column, restores the complete range of the table cell, and aggregates the discrete rotation detection box into a structured text unit with row and column indexes. It fully preserves the nested, multi-column and multi-cell layout structure of the bill table, and realizes the hierarchical integration of detection fragments into structured cell units.
[0128] S2565: Binarize the text unit image, extract the pixels of each column in the binarized text unit image, calculate the sum of the pixels, and use it as the vertical projection accumulation value of the pixel column (that is, starting from the leftmost column of the text unit image, process each pixel column in turn, each column contains all the pixels in the column from top to bottom (the number is equal to the height of the text unit image), take out all the pixels in the current column, and accumulate the gray values of each pixel (in the binarized text unit image, the foreground pixels are usually white / high gray values, and the background is black / low gray values), and the accumulation result is the vertical projection accumulation value of the pixel column).
[0129] The vertical projection summation values of all pixel columns are sorted in ascending order according to the column index of the pixel (the column index is the sequence number of the pixel column (0, 1, 2, 3...)) to form a vertical projection curve (that is, the vertical projection curve is plotted with the column index as the horizontal axis and the vertical projection summation value as the vertical axis; this vertical projection curve reflects the grayscale distribution of each pixel column, with the column containing character strokes having a higher projection value and the column containing character gaps having a lower projection value).
[0130] A Gaussian kernel window of a fixed size is preset for the vertical projection curve. A sliding convolution is performed along the column index direction of the vertical projection curve. A weighted average is calculated for the accumulated vertical projection values of each pixel column within the Gaussian kernel window. This weighted average is used as the new accumulated vertical projection value for the center pixel column of the Gaussian kernel window, resulting in a smooth vertical projection curve. (That is, a Gaussian kernel window of a fixed size is preset (e.g., window size 5 or 7), the weight values within the Gaussian kernel window conform to a Gaussian distribution, with the largest weight in the middle and gradually decreasing towards both ends; the Gaussian kernel window is moved along the column index direction of the vertical projection curve from left to right...) Slide to the right. At each Gaussian kernel window position, calculate the weighted average of the vertical projection accumulation value of each pixel column within the Gaussian kernel window and the corresponding Gaussian weight. Replace the vertical projection accumulation value at the center column position of the window with this weighted average value (that is, replace the vertical projection accumulation value of the pixel column at the center position within the Gaussian kernel window). During the sliding process, the window center moves column by column until all columns have been processed. After Gaussian filtering, the spikes and noise in the vertical projection curve are eliminated, and the vertical projection curve becomes smoother, while retaining the overall trend of peaks and troughs, resulting in a smooth vertical projection curve.
[0131] Whether the cumulative vertical projection value of the selected pixel within the smooth vertical projection curve is less than the cumulative vertical projection value of the pixel's preceding and following pixels;
[0132] If so, then that pixel is taken as a local minimum point;
[0133] Calculate the average value of the accumulated vertical projection values of the smooth vertical projection curve as the global average projection value;
[0134] Calculate the ratio between the global average projection value and the cumulative vertical projection value of the local minimum point; if this ratio is less than a preset projection ratio threshold (i.e., preferably 0.4 ~ 0.6, preferably 0.5 in this embodiment (i.e., when the cumulative projection value of the local minimum point is less than 0.5 times the global average projection value, it is determined to be a valid valley)), then the pixel column corresponding to the local minimum point is determined as a valid vertical valley column, and the column indices of all valid vertical valley columns are used as candidate vertical segmentation points (i.e., scan each pixel on the smooth vertical projection curve; the projection value of this pixel is less than the projection value of its left adjacent pixel and less than the projection value of its right adjacent pixel, and the pixel that satisfies both of these conditions is a local minimum point (valley position). The process involves: calculating the projection values of all pixels on the smoothed vertical projection curve, averaging the cumulative vertical projection values of the smoothed vertical projection curve, and using this average as the global average projection value; calculating the ratio of the cumulative vertical projection value to the global average projection value for each local minimum point; and if this ratio is less than a preset projection ratio threshold (meaning the trough is low enough and represents a true character gap), then determining the column index corresponding to the local minimum point as a valid vertical trough column, collecting the column indices of all valid vertical trough columns, and identifying the vertical separation positions between characters.
[0135] The sum of gray values of the pixels in each row of the binarized text unit image is calculated and used as the horizontal projection accumulation value (that is, starting from the top row of the image, each row of pixels is processed in turn, each row contains all the pixels from left to right in that row (the number is equal to the image width), all the pixels in the current row are taken out, and the gray values of each pixel are summed up. The summation result is the horizontal projection accumulation value of that pixel row).
[0136] Arrange the accumulated horizontal projection values of all pixel rows in ascending order of row index to form a horizontal projection curve; filter the pixel rows in the horizontal projection curve with a accumulated horizontal projection value of zero, and take the consecutive pixel rows as the consecutive row index interval (that is, the consecutive interval with a projection value of zero, which comes from the "pure background blank line" between the character lines in the binarized image, that is, the white paper part without ink).
[0137] The row of pixels located in the middle position in the continuous row index interval is taken as the candidate horizontal segmentation point;
[0138] It should be noted that in the above steps, the vertical and horizontal projection curves are calculated for the binarized text unit image. The vertical projection is smoothed by Gaussian sliding window to eliminate projection noise. Reliable vertical character segmentation valleys are extracted by local minimum value screening and global projection ratio constraint. The horizontal projection directly identifies blank intervals without characters and takes the midpoint of the interval as the horizontal segmentation benchmark. The dual projection mechanism captures the left and right gaps and the blanks between the upper and lower lines of the characters, filters out false segmentation points formed by stains and stroke noise, accurately locks the natural horizontal and vertical segmentation positions of the characters, and generates highly reliable horizontal and vertical segmentation candidate points, providing a basis for accurate character segmentation.
[0139] S2566: Vertically segment the text unit image according to the pixel column corresponding to the candidate vertical segmentation point, and horizontally segment it according to the pixel row corresponding to the candidate horizontal segmentation point to obtain multiple candidate character blocks;
[0140] For each candidate character block, the number of foreground pixels is calculated to the total number of pixels in the candidate character block, and the ratio is further calculated as the foreground pixel density.
[0141] If the foreground pixel density is greater than or equal to a preset density threshold, then the candidate character block is taken as the target character block;
[0142] Filter all foreground pixels of the target character block, and extract the bounding rectangle boundary of all foreground pixels (that is, for each target character block, scan all foreground pixels inside it, find the minimum row index, maximum row index, minimum column index, and maximum column index of all foreground pixels, and the four boundary values determine the rectangular area that encloses all foreground pixels).
[0143] The outer rectangular boundary of the corresponding position is cropped inside the target character block to obtain the initial character image (with the extra blank borders removed); the initial character image is input to a preset uniform pixel size (i.e., 32×32 or 64×64 pixels, keeping the character ratio unchanged during scaling, and filling the insufficient parts) to obtain a standardized character image;
[0144] Extract all foreground pixels in the standardized character image, and perform four-neighbor connectivity labeling on the foreground pixels in the standardized character image (i.e., scan each foreground pixel in the image, check whether the pixels adjacent to it in the four directions of up, down, left, and right are also foreground pixels, and classify the interconnected foreground pixels into the same verification connectivity region, and assign a unique label number to each connectivity region), to obtain multiple foreground region connectivity regions.
[0145] If the number of connected components in the foreground region is greater than 1, then calculate the area of each connected component in the foreground region, calculate the area ratio of the area of the connected component in the foreground region to the total area of all connected components in the foreground region, and select the connected component in the foreground region corresponding to the area ratio of the maximum value as the main connected component.
[0146] The foreground pixels of the remaining foreground region connected regions outside the main connected region are assigned as background pixels to obtain a single connected region character image;
[0147] Collect all the single-connected character images corresponding to the target character blocks as a single character image;
[0148] It should be noted that in the above steps, candidate character blocks are obtained by segmenting the image using candidate horizontal and vertical segmentation points. Blank noise blocks are filtered by using a foreground pixel density threshold (preferably 0.05 ~ 0.2, preferably 0.1 in this embodiment). The bounding rectangle of the foreground pixels is extracted to remove redundant blank borders, and the images are scaled and filled proportionally to generate standardized character images of uniform size. Four-neighbor connected component detection is used to identify redundant noise points and separate small impurities within the characters. The main connected component with the highest area ratio is retained and the remaining discrete foreground pixels are removed. The final output is a single character image containing only a single complete character without any redundant impurities. The entire process eliminates character segmentation defects caused by ticket tilting, text adhesion, printing stains, and detection fragments, and outputs standardized single-character materials with uniform size, clean background, and no noise interference.
[0149] Example 2
[0150] like Figure 7 As shown, the present invention also provides a YOLOv8-based ticket document rotation target information detection and positioning system, including: a data acquisition module 10 and an analysis module 20;
[0151] The acquisition module 10 is used to acquire the image of the invoice document to be detected, and to normalize the size of the image of the invoice document to be detected to obtain a standardized input tensor.
[0152] The analysis module 20 is used to input the normalized input tensor into the backbone network of YOLOv8 and output a multi-layer feature map set; to analyze the multi-layer feature map set according to spatial resolution and construct an enhanced feature map pyramid; to set multiple sizes of rotation anchor boxes for the enhanced feature map pyramid and input them into the detection head of YOLOv8 to output target rotation bounding boxes; and to analyze the target rotation bounding boxes to obtain a single character image.
[0153] The aforementioned acquisition module performs standardized preprocessing and size normalization on the original images of the tickets to be detected. It uses proportional scaling combined with edge filling to eliminate size distortion problems of ticket images from different shooting angles, resolutions, and sizes, fully preserving all graphic and textual information of the ticket text, tables, and lines. At the same time, it completes pixel value normalization and image dimension format conversion, unifying the data distribution and tensor dimension specifications of the model input, eliminating model inference disturbances caused by differences in brightness, inconsistent sizes, and chaotic formats of the original images, and providing stable, standardized, and distortion-free input tensors for subsequent network feature extraction, ensuring the stability and robustness of subsequent feature extraction and object detection.
[0154] The analysis module relies on the YOLOv8 backbone network to complete multi-scale deep feature extraction. It constructs an enhanced feature map pyramid by combining spatial pyramid pooling and bidirectional deep and shallow feature fusion strategies, enabling each level of features to simultaneously possess shallow character detail texture and global layout semantic information, taking into account the recognition needs of small character details and large-size table layouts. By configuring multi-size, multi-aspect-ratio, and multi-angle rotating anchor boxes, it breaks through the limitation of traditional horizontal anchor boxes that only adapt to horizontal text, accurately adapting to complex conditions such as slanted text, deflected cells, and diagonal layout on invoices, and outputting rotating target bounding boxes that fit the true shape of the text. At the same time, it combines layout topology analysis, adaptive spacing determination, horizontal and vertical projection segmentation, and connected component denoising mechanisms to complete invoice field aggregation and refined single-character segmentation, effectively improving problems such as character adhesion, noise interference, and irregular segmentation. This module realizes integrated intelligent processing of invoice rotating target detection, layout structure analysis, field unit division, and accurate single-character segmentation, significantly improving the accuracy of complex invoice text localization and character segmentation, and effectively making up for the shortcomings of traditional detection systems in terms of low intelligence and weak scene adaptability.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; those skilled in the art can modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting and locating rotational target information in YOLOv8 invoice documents, characterized in that, Specific operating methods include: Obtain the image of the invoice document to be detected, and normalize the size of the image to obtain a standardized input tensor; The standardized input tensor is input into the backbone network of YOLOv8, and a multi-layer feature map set is output. The multi-layer feature map set is analyzed according to spatial resolution to construct an enhanced feature map pyramid. Multiple sizes of rotational anchor boxes are set for the enhanced feature map pyramid and input into the detector head of YOLOv8 to output the target rotational bounding box. The target rotational bounding box is analyzed to obtain a single character image.
2. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 1, characterized in that, The standardized input tensor is input into the YOLOv8 backbone network to output a multi-layer feature map set. The specific operation steps are as follows: The normalized input tensor is input into the backbone network of YOLOv8. The normalized input tensor is then passed forward through multiple sequentially stacked convolutional blocks and multiple downsampling layers within the backbone network. The normalized input tensor is downsampled according to the multiple downsampling layers, and the intermediate feature map corresponding to the downsampling layer is output. The spatial resolution and number of channels of each intermediate feature map are recorded. Obtain the intermediate feature map output by the last downsampling layer in the backbone network, and use it as the highest layer feature map; A spatial pyramid operation is performed on the highest-level feature map. Parallel pooling operations are then performed on the highest-level feature map using multiple pooling windows of increasing size within the spatial pyramid to obtain multiple pooling results. The multiple pooling results are then concatenated according to the number of channels to obtain a multi-scale pooled feature map. Intermediate feature maps output by each downsampling layer in the backbone network, except for the last downsampling layer, are then sequentially obtained. The multi-scale pooling feature map is upsampled to the same spatial resolution as each intermediate feature map. The upsampled multi-scale pooling feature map is then fused with each intermediate feature map element-wise to obtain multiple fused feature maps. A convolution transformation is performed on each fused feature map to obtain the hierarchical feature map corresponding to the downsampled layer. All hierarchical feature maps corresponding to the downsampled layers are combined with the highest layer feature map to form a multi-layer feature map set.
3. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 2, characterized in that, An enhanced feature map pyramid is constructed by analyzing the multi-layer feature map set according to spatial resolution; multiple sizes of rotating anchor boxes are set for the enhanced feature map pyramid and input into the YOLOv8 detection head to output the target rotating bounding box; the target rotating bounding box is analyzed to obtain a single character image. The specific operation steps are as follows: The highest-level feature map in the multi-layer feature map set is selected as the first reference feature map; the fusion feature map with the second smallest spatial resolution in the multi-layer feature map set is selected as the first mesoscale layer feature map. The fused feature map with the highest spatial resolution is selected as the feature map of the bottom scale layer; The first reference feature map is subjected to a first upsampling operation according to the second smallest spatial resolution of the first mesoscale layer feature map to obtain the upsampled first reference feature map; The upsampled first baseline feature map and the first mesoscale layer feature map are concatenated along the number of channels to obtain a first concatenated feature map; the first concatenated feature map is subjected to a first convolutional transformation to obtain a first intermediate fused feature map; the first intermediate fused feature map is subjected to a second upsampling operation, and the spatial resolution of the first intermediate fused feature map after the second upsampling is adjusted according to the maximum spatial resolution of the bottom-scale layer feature map; the first intermediate fused feature map after the second upsampling is concatenated with the bottom-level feature map along the number of channels to obtain a second concatenated feature map; A second convolutional transformation is performed on the second concatenated feature map to obtain a second intermediate fused feature map, which serves as the first enhanced output feature map. A first downsampling operation is then performed on the first enhanced output feature map, adjusting its spatial resolution according to the spatial resolution of the first intermediate fused feature map. The first enhanced output feature map and the first intermediate fused feature map are then concatenated along the number of channels to obtain a third concatenated feature map. A third convolutional transformation is performed on the third concatenated feature map to obtain a third intermediate fused feature map, which serves as the second enhanced output feature map. A second downsampling operation is performed on the second enhanced output feature map. The spatial resolution of the second enhanced output feature map after the second downsampling is adjusted according to the spatial resolution of the first reference feature map. The second enhanced output feature map after the second downsampling is concatenated with the first reference feature map along the number of channels to obtain a fourth concatenated feature map. A fourth convolution transformation is performed on the fourth concatenated feature map to obtain a fourth intermediate fusion feature map, which is used as the third enhanced output feature map. The first enhanced output feature map, the second enhanced output feature map, and the third enhanced output feature map are stacked in ascending order of spatial resolution to form a three-layer enhanced feature map pyramid.
4. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 3, characterized in that, Multiple rotating bounding boxes of various sizes are set for the enhanced feature map pyramid and input into the detection head of YOLOv8 to output the target rotating bounding box; the target rotating bounding box is analyzed to obtain a single character image. The specific operation steps are as follows: Each layer of the enhanced output feature map in the enhanced feature map pyramid is uniformly divided into grid cells, and multiple rotating anchor frames of various specifications are set according to the anchor point position of the grid cells. Multiple sizes of rotating anchor frames are input into the regression branch of the detection head in YOLOv8 to obtain the predicted rotating bounding box; Furthermore, the detection head uses a classification branch containing preset categories to analyze each rotating anchor frame simultaneously to obtain the target rotating bounding box; it rotates adjacent target rotating bounding boxes and calculates the Euclidean distance between their center points; it uses the Euclidean distance between their center points to filter adjacent bounding box pairs; it determines the connection relationship type of the adjacent bounding boxes and constructs a layout structure topology diagram of the document. The connected component algorithm is used to analyze the topology of the layout structure to obtain individual character images.
5. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 4, characterized in that, Each layer of the enhanced output feature map in the enhanced feature map pyramid is uniformly divided into grid cells, and multiple rotating anchor frames of various specifications are set according to the anchor point position of the grid cells. Multiple rotating anchor frames of various sizes are input into the regression branch of the detection head in YOLOv8 to obtain predicted rotating bounding boxes. Then, the classification branch within the detection head, containing preset categories, is used to analyze each rotating anchor frame simultaneously to obtain the target rotating bounding box. The specific operation steps are as follows: For each layer of enhanced output feature map in the enhanced feature map pyramid, the enhanced output feature map of that layer is evenly divided into grid cells, and each grid cell corresponds to an anchor point position; multiple sizes of rotating anchor frames are set using the x-coordinate of the center point, y-coordinate of the center point, width parameter, height parameter, and preset initial rotation angle parameter of each anchor point position; Based on each rotated anchor box at the anchor point position in the enhanced output feature map of each layer, extract the feature vector of the anchor point position corresponding to each rotated anchor box. The feature vector of the anchor point position corresponding to each rotating anchor frame is input into the regression branch of the YOLOv8 detector head. The regression branch outputs the x-coordinate offset, y-coordinate offset, width scaling, height scaling, and rotation angle offset of the rotating anchor frame's center point. The x-coordinate offset is summed with the x-coordinate of the rotating anchor frame's center point to obtain the predicted x-coordinate. The y-coordinate offset is summed with the y-coordinate of the rotating anchor frame's center point to obtain the predicted y-coordinate. The width scaling is multiplied by the width parameter of the rotating anchor frame to obtain the predicted width. The height scaling is multiplied by the height parameter of the rotating anchor frame to obtain the predicted height. The rotation angle offset is summed with the rotation angle parameter of the rotating anchor frame to obtain the predicted rotation angle. A predicted rotating bounding box is constructed based on the predicted x-coordinate, y-coordinate, width, height, and rotation angle. The feature vectors at the anchor points corresponding to each rotating anchor frame are analyzed simultaneously using the classification branches containing preset categories within the detection head, and the confidence score vectors of each preset category corresponding to the rotating anchor frame are output; the confidence score vector with the maximum value is selected as the predicted confidence score of the rotating anchor frame, and the preset category corresponding to the predicted confidence score is selected as the predicted category of the rotating anchor frame. And extract the predicted rotated bounding box corresponding to the predicted confidence score; If the predicted confidence score is greater than or equal to the preset confidence threshold, then the predicted rotated bounding box is taken as a candidate rotated bounding box. Calculate the rotation intersection-union ratio between any two candidate rotated bounding boxes; if the rotation intersection-union ratio is greater than a preset intersection-union ratio threshold, then select the candidate rotated bounding box with the maximum predicted confidence score between the two candidate rotated bounding boxes. Iterate through all the selected candidate rotated bounding boxes until the rotation intersection-union ratio (CUI) between any two selected candidate rotated bounding boxes is less than or equal to the CUI threshold. Finally, retain all the retained candidate rotated bounding boxes as the target rotated bounding boxes.
6. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 5, characterized in that, The process involves rotating adjacent target bounding boxes and calculating the Euclidean distance between their center points; using this Euclidean distance to filter adjacent bounding box pairs; determining the connection type of the adjacent bounding boxes; and constructing a layout topology diagram of the document. The specific steps are as follows: The orientation of the target rotating bounding box in the image of the document to be detected is determined based on the predicted rotation angle of the target rotating bounding box; the spatial positioning of the target rotating bounding box is determined based on the predicted x-coordinate and y-coordinate of the predicted center point of the target rotating bounding box. Using the spatial location of the target rotating bounding box as the rotation center, a rotation transformation is performed according to the negative angle value of the predicted rotation angle, so that the orientation of the target rotating bounding box after the rotation transformation is corrected to the horizontal direction. A rectangular local region is cropped from the enhanced output feature map after the rotation transformation according to the predicted width and predicted height. A pooling operation is performed on the cropped rectangular local region, and the local region feature map is unfolded into a one-dimensional feature vector, which is used as the local region feature vector corresponding to the target rotating bounding box. Based on the predicted x-coordinate and y-coordinate of the center point of all target rotated bounding boxes, the Euclidean distance between the center points of every two target rotated bounding boxes is calculated in a two-dimensional plane coordinate system. If the Euclidean distance is less than a preset distance threshold, the two target rotated bounding boxes are paired as adjacent bounding boxes. All adjacent bounding box pairs are obtained by traversing all rotating bounding box combinations. The first local region feature vector corresponding to the first target rotated bounding box and the second local region feature vector corresponding to the second target rotated bounding box in the adjacent bounding box pair are obtained. The difference between the first local region feature vector and the second local region feature vector is calculated as the lookup vector. The first local region feature vector and the second local region feature vector are summed to obtain a sum vector. The difference vector and the sum vector are concatenated according to the channel dimension to obtain a concatenation relationship feature vector. The concatenation relationship feature vector is input into a pre-trained relationship prediction network, and the connection relationship type is output. Each target rotated bounding box is used as a topology node, and the layout structure topology of the ticket document is constructed using the connection relationship type of all adjacent bounding box pairs as the edge attributes between the corresponding topology nodes.
7. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 6, characterized in that, The connected component analysis algorithm is used to analyze the topology of the page layout to obtain individual character images. The specific operation steps are as follows: Based on the analysis of the predicted center point x-coordinate and predicted center point y-coordinate of the target rotated bounding box in the layout structure topology diagram, row groups of topology nodes are obtained; the horizontal projection interval is determined by obtaining the left and right boundary x-coordinate values of the target rotated bounding boxes in the row groups; a horizontal overlap relationship diagram is constructed for the horizontal projection intervals of any two target rotated bounding boxes in the row groups; a connected component algorithm is performed on the horizontal overlap relationship diagram to obtain an ordered bounding box sequence within the row; adjacent target rotated bounding boxes in the ordered bounding box sequence within the row are extracted and analyzed to obtain field unit groups; a text unit image of the document is constructed from multiple field unit groups; the text unit image is binarized to extract candidate vertical and candidate horizontal segmentation points; the text unit image is segmented using the candidate vertical and candidate horizontal segmentation points to obtain target character blocks; the circumscribed rectangle boundary of the foreground pixels of the target character blocks is extracted, and the circumscribed rectangle boundary of the corresponding position is cropped inside the target character block to obtain a standardized character image; a connected component analysis is performed on the standardized character image to obtain a single character image.
8. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 7, characterized in that, Based on the analysis of the predicted center point x-coordinate and predicted center point y-coordinate of the target rotated bounding box in the layout structure topology diagram, row groups of topology nodes are obtained; the horizontal projection interval is determined by obtaining the left and right boundary x-coordinate values of the target rotated bounding box in the row group; a horizontal overlap relationship graph is constructed for the horizontal projection intervals of any two target rotated bounding boxes in the row group; the connected component algorithm is performed on the horizontal overlap relationship graph to obtain the ordered bounding box sequence within the row. The specific operation steps are as follows: Obtain the x-coordinate and y-coordinate of the predicted center point of the target rotated bounding box of all topological nodes in the layout structure topology diagram. Sort all topological nodes in ascending order according to the value of the predicted center point y-coordinate to obtain the y-coordinate sorted node sequence. Select the first topological node in the y-coordinate sorted node sequence as the starting node of the current row group, and mark the predicted center point y-coordinate of the starting node as the reference y-coordinate of the current row. Iterate through all the topological nodes in the sorted node sequence by vertical coordinate, and calculate the difference between the vertical coordinate of the predicted center point of the currently traversed topological node and the vertical coordinate of the current row reference vertical coordinate. If the difference in the ordinate is less than the preset row height threshold, the current topology node is assigned to the current row group, and the current row reference ordinate is updated to the average of the predicted center point ordinates of all topology nodes in the current row group. If the difference in the ordinate is greater than or equal to the row height threshold, then the current row group is closed, and a new row group is created with the current topology node as the starting node of the new row group. The ordinate of the predicted center point of the starting node of the new row group is marked as the reference ordinate of the new row. The topology nodes are traversed until all topology nodes in the ordinate sorting node sequence have been traversed, resulting in multiple row groups. Get the left and right x-coordinates of the target rotated bounding box for all topological nodes in the current row group; The horizontal projection interval of each target rotated bounding box is determined based on the left and right boundary x-coordinate values; the overlap length of the horizontal projection intervals of any two target rotated bounding boxes within a row group is calculated; and the predicted width of the minimum value between these two target rotated bounding boxes is selected. Calculate the overlap ratio between the overlap length and the predicted width of the minimum value; if the overlap ratio is greater than the preset overlap ratio threshold, determine that there is a horizontal overlap relationship between any two target rotating bounding boxes, and record the two target rotating bounding boxes as a horizontal overlap pair; construct a horizontal overlap relationship map within the row group based on all horizontal overlap pairs; Perform connected component analysis on the horizontal overlap graph to merge the target rotated bounding boxes of horizontal overlap into the same overlap group, resulting in multiple overlap groups within that row group; for the target rotated bounding boxes within each overlap group, arrange them in ascending order according to the left boundary x-coordinate value to obtain an ordered subsequence within the overlap group; For the target rotated bounding boxes within the row group that are not covered by the horizontal overlapping pairs, they are taken as isolated rotated bounding boxes; they are sorted in ascending order according to the left boundary x-coordinate values to obtain independent subsequences; the ordered subsequences within the overlapping groups and the independent subsequences are merged and sorted according to their respective left boundary x-coordinate values to obtain the in-row ordered bounding box sequence corresponding to the row group.
9. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 8, characterized in that, Extracting adjacent target bounding boxes within an inline ordered bounding box sequence and analyzing them to obtain field cell groups; constructing a text cell image of the document from multiple field cell groups. The specific steps are as follows: Traverse the ordered bounding box sequence within each row group, extract the target rotation bounding box located on the left of two adjacent target rotation bounding boxes as the left bounding box, and the target rotation bounding box located on the right as the right bounding box; extract the right boundary x-coordinate value of the left bounding box as the right edge position of the left bounding box; extract the left boundary x-coordinate value of the right bounding box as the left edge position of the right bounding box; calculate the absolute value of the difference between the right edge position of the left bounding box and the left edge position of the right bounding box as the horizontal spacing between two adjacent target rotation bounding boxes; The predicted width of the left bounding box is used as the width of the left bounding box, and the predicted width of the right bounding box is used as the width of the right bounding box. Calculate the ratio of the horizontal spacing to the width of the left bounding box, and use it as the left spacing ratio; Calculate the ratio of the horizontal spacing to the width of the right bounding box, and use it as the right spacing ratio; The ratio of the minimum value selected from the left spacing ratio and the right spacing ratio is taken as the effective spacing ratio between two adjacent target rotated bounding boxes; A preset first spacing threshold and a second spacing threshold are used; it is then determined whether the effective spacing ratio is less than the first spacing threshold. If yes, then the two adjacent target rotated bounding boxes are marked as internal links, indicating that they belong to the same field unit within the same cell; if no, then it is determined whether the effective spacing ratio is less than the second spacing threshold; if yes, then the two adjacent target rotated bounding boxes are marked as field spacing; if no, then the two adjacent target rotated bounding boxes are marked as column spacing. Merge adjacent target rotated bounding boxes marked as internally connected within the ordered bounding box sequence to obtain an initial character fragment; use the field interval and column interval as mandatory delimiters to divide the ordered bounding box sequence into multiple consecutive fragment groups, which serve as multiple field unit groups; Get the left boundary x-coordinate value of all target rotated bounding boxes in each field cell group, filter the minimum left boundary x-coordinate value as the column start boundary of the field cell group; get the right boundary x-coordinate value of all target rotated bounding boxes in each field cell group, filter the maximum right boundary x-coordinate value as the column end boundary of the field cell group. The horizontal span interval of each field unit group is determined based on the column start boundary and the column end boundary; the difference between the column start boundary and the column end boundary of any two field unit groups is calculated; if the difference between the column start boundary and the difference between the column end boundary are less than a preset alignment threshold, then the two field units are combined into the same column group to obtain the text unit image of the invoice document.
10. The method for detecting and locating rotating target information in YOLOv8 invoice documents according to claim 9, characterized in that, The text unit image is subjected to binarization analysis to extract candidate vertical and horizontal segmentation points; the text unit image is segmented using the candidate vertical and horizontal segmentation points to obtain target character blocks; the bounding rectangle boundaries of the foreground pixels of the target character blocks are extracted, and the bounding rectangle boundaries of the corresponding positions are cropped inside the target character blocks to obtain standardized character images; the standardized character images are subjected to connected component analysis to obtain individual character images. The specific operation steps are as follows: The text unit image is binarized, and the pixels in each column of the binarized text unit image are extracted. The sum of the pixels is calculated as the vertical projection accumulation value of that column. The vertical projection accumulation values of all pixel columns are arranged in ascending order according to the column index of the pixels to form a vertical projection curve. A Gaussian kernel window of a preset fixed size is applied to the vertical projection curve, and a sliding convolution is performed along the column index direction of the vertical projection curve. The weighted average value of the vertical projection accumulation value of each pixel column within the Gaussian kernel window is calculated, and this weighted average value is used as the new vertical projection accumulation value of the center pixel column of the Gaussian kernel window to obtain a smooth vertical projection curve. If the cumulative vertical projection value of the selected pixel within the smooth vertical projection curve is less than the cumulative vertical projection value of the pixel's preceding and following pixels, then the pixel is considered a local minimum. The average value of the cumulative vertical projection values of the smooth vertical projection curve is calculated as the global average projection value. The ratio between the global average projection value and the cumulative vertical projection value of the local minimum is calculated. If the ratio is less than the preset projection ratio threshold, the pixel column corresponding to the local minimum point is determined as a valid vertical valley column, and the column index of all valid vertical valley columns is used as a candidate vertical segmentation point; the sum of gray values of the pixels in each row of the binarized text unit image is calculated as the horizontal projection accumulation value. The horizontal projection summation values of all pixel rows are sorted in ascending order by row index to form a horizontal projection curve. Pixel rows with a horizontal projection summation value of zero are selected from the horizontal projection curve, and consecutive pixel rows are designated as consecutive row index intervals. The pixel row located in the middle of the consecutive row index interval is selected as a candidate horizontal segmentation point. The text unit image is vertically segmented according to the pixel column corresponding to the candidate vertical segmentation point, and horizontally segmented according to the pixel row corresponding to the candidate horizontal segmentation point, resulting in multiple candidate character blocks. For each candidate character block, the number of foreground pixels is calculated relative to the total number of pixels within the candidate character block, and the ratio is further calculated as the foreground pixel density. If the foreground pixel density is greater than or equal to a preset density threshold, then the candidate character block is taken as the target character block; Filter all foreground pixels of the target character block and extract the bounding rectangle boundaries of all foreground pixels; crop the bounding rectangle boundaries of the corresponding positions inside the target character block to obtain an initial character image; input the initial character image into a preset uniform pixel size to obtain a standardized character image; Extract all foreground pixels from the standardized character image, and perform four-neighbor connectivity labeling on the foreground pixels in the standardized character image to obtain multiple foreground region connected components; If the number of connected components in the foreground region is greater than 1, then calculate the area of each connected component in the foreground region, calculate the area ratio of the area of the connected component in the foreground region to the total area of all connected components in the foreground region, and select the connected component in the foreground region corresponding to the area ratio of the maximum value as the main connected component. The foreground pixels of the remaining foreground region connected regions outside the main connected region are assigned as background pixels to obtain a single connected region character image; Collect all the single-connected character images corresponding to the target character blocks as a single character image.