Image orientation recognition method, image orientation recognition model, and storage medium

The image orientation identification method improves accuracy by detecting character regions and fusing positional and morphological features with text line classification results, addressing the limitations of conventional deep convolutional networks.

JP7856174B2Active Publication Date: 2026-05-11RICOH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
RICOH CO LTD
Filing Date
2025-01-17
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Conventional image direction identification methods using deep convolutional networks suffer from low accuracy due to the inability to effectively utilize character features, leading to reduced classification accuracy.

Method used

An image orientation identification method that detects character regions, extracts positional and morphological features, and fuses them with text line classification results to improve orientation recognition, using pre-trained modules for character region detection and feature fusion.

Benefits of technology

Enhances image orientation identification accuracy by reducing noise and improving performance through the integration of positional and morphological features, while reducing the need for extensive training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007856174000001
    Figure 0007856174000001
  • Figure 0007856174000002
    Figure 0007856174000002
  • Figure 0007856174000003
    Figure 0007856174000003
Patent Text Reader

Abstract

To provide a method and model for identifying an image orientation, and a storage medium.SOLUTION: An image orientation identification method includes: detecting character regions from a target image to be identified, and obtaining position coordinates and region features for each character region; identifying image features in the target image of the character region, based on the position coordinates, to be merged with the region features, to generate first merge features of the character regions, and generating text row classification results of the character regions based on the first merge features; acquiring position features and form features for each character region, performing merge for the text row classification results based on the position features and the form features, and generating an orientation identification result of the target image. The invention improves accuracy of identifying image orientation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and specifically to a method for image direction identification, an image direction identification model (also referred to as an image direction identification device), and a storage medium.

Background Art

[0002] In intelligent office operations, it is common to perform character identification, such as OCR (Optical Character Recognition), on document images scanned and uploaded by users. At that time, in order to ensure the accuracy of the identification result, it is necessary to ensure that the image direction is the correct direction before performing the identification. However, due to various factors, the direction of the image generated by the user's scan may be 0°, 90°, 180°, 270° (where 0° is the correct direction), and the image direction is not limited to the correct direction. Therefore, it is necessary to rotate the image manually or by an algorithm to make it in the correct direction before performing the identification. Here, it is assumed that the character direction is parallel to the long side or the short side of the image, and the slight inclination of the character direction with respect to the length / width direction of the image is not considered.

[0003] The purpose of image direction identification is to rotate the image in the correct direction. Currently, as a main method of image direction identification, there is a method of directly classifying an image through a depth (deep) convolutional network. However, the conventional depth convolutional network extracts features for the entire image, and the character part cannot represent an important role therein, and by introducing a lot of noise, the classification accuracy is reduced, and the accuracy of image direction identification is low. Therefore, a method for improving the accuracy of image direction identification is needed.

Summary of the Invention

Problems to be Solved by the Invention

[0004] At least one embodiment of the present invention provides an image direction identification method, a model, and a storage medium that can solve the problem of low accuracy of image direction identification in the prior art. [Means for solving the problem]

[0005] To solve the above problems, first, the first embodiment of the present invention is: The system detects character regions from the target image to be identified, and obtains the position coordinates and region features of each character region; Based on the position coordinates, the image features in the target image of the character region are identified and merged with the region features to generate a first merged feature of the character region; and based on the first merged feature, the text line classification result of the character region is generated; The present invention provides an image orientation identification method that includes acquiring positional features and morphological features for each character region, fusing the positional features and morphological features with the text line classification results, and generating the orientation identification result for the target image.

[0006] As an option, positional and morphological features are acquired for each character region, and based on the positional and morphological features, the text line classification results are merged to generate the direction identification result for the target image. The process includes inputting the position coordinates of each character area and the text line classification result into an image orientation identification module, and obtaining the orientation identification result of the target image output by the image orientation identification module, The aforementioned image orientation identification module is For each character region, the positional features and morphological features are identified based on the positional coordinates, the positional features and morphological features are merged to generate a second fused feature, and the weights of the character region are generated based on the second fused feature; Based on the weight of each character region, the text line classification results for that character region are merged to generate the direction identification result for the target image.

[0007] Optionally, the positional features include the relative pole angle, which is the difference between the pole angle and the reference pole angle, and the pole diameter. The aforementioned morphological features include at least one of the height, width, area, contour length, and center of mass of the character region.

[0008] As an option, it is possible to detect character regions from the target image to be identified and obtain the position coordinates and region features for each character region. This includes inputting the target image into a character region detection module and obtaining the position coordinates and region features of each character region output by the character region detection module.

[0009] As an option, based on the position coordinates, the image features in the target image of the character region are identified and merged with the region features to generate a first merged feature of the character region, and based on the first merged feature, the text line classification result of the character region is generated. This includes inputting the target image and the position coordinates and region features of the character region into a feature fusion and text line classification module, and generating a text line classification result for the character region output by the feature fusion and text line classification module.

[0010] Optionally, the image orientation identification method, before detecting the character region from the target image, The character region detection module is pre-trained using first training data, which includes multiple first images on which character regions are marked; Second training data, including multiple second images on which the text line classification results are indicated, is input to the trained character region detection module; the feature fusion and text line classification module is pre-trained using the position coordinates and region features of the character regions output by the character region detection module, and the second training data; A third training data set is obtained, which includes multiple third images in the target region with image orientations indicated; The method includes inputting the third training data into the trained character region detection module, inputting the position coordinates and region features of the character regions output by the character region detection module, and the third training data into a trained feature fusion and text line classification module, and training the image orientation discrimination module using the text line classification results output by the feature fusion and text line classification module and the third training data.

[0011] Furthermore, a second embodiment of the present invention is: A character region detection module that detects character regions from an image to be identified and obtains the position coordinates and region features of each character region, A feature fusion and text line classification module that, based on the position coordinates, identifies image features in the target image of the character region and fuses them with the region features to generate a first fused feature of the character region, and generates a text line classification result of the character region based on the first fused feature, The present invention provides an image orientation recognition model that includes an image orientation recognition module that acquires positional features and morphological features for each character region, fuses the positional features and morphological features with the text line classification results, and generates an image orientation recognition result for the target image.

[0012] As an option, the image orientation identification module further: For each character region, the positional features and morphological features are identified based on the positional coordinates, the positional features and morphological features of the character region are merged to generate a second fused feature, and the weights of the character region are generated based on the second fused feature; Based on the weight of each character region, the text line classification results for that character region are merged to generate the direction identification result for the target image.

[0013] Optionally, the positional features include the relative pole angle, which is the difference between the pole angle and the reference pole angle, and the pole diameter. The aforementioned morphological features include at least one of the height, width, area, contour length, and center of mass of the character region.

[0014] Optionally, the character region detection module further generates an image feature diagram corresponding to the target image, detects the character region based on the image feature diagram, obtains the position coordinates of the character region, and extracts features corresponding to the character region from the image feature diagram to define the region features of the character region.

[0015] As an option, the image orientation recognition model is The character region detection module is pre-trained using first training data, which includes multiple first images on which character regions are marked; Second training data, including multiple second images on which the text line classification results are indicated, is input to the trained character region detection module; the feature fusion and text line classification module is pre-trained using the position coordinates and region features of the character regions output by the character region detection module, and the second training data; A third training data set is obtained, which includes multiple third images in the target region with image orientations indicated; The present invention further includes a training module that inputs the third training data into the trained character region detection module, inputs the position coordinates and region features of the character regions output by the character region detection module, and the third training data into a trained feature fusion and text line classification module, and trains the image orientation discrimination module using the text line classification results output by the feature fusion and text line classification module and the third training data.

[0016] Furthermore, a third embodiment of the present invention provides a computer-readable storage medium in which a cocomputer program is stored, and the above-described image orientation identification method is realized by causing a processor to execute the computer program.

[0017] Furthermore, a fourth embodiment of the present invention provides a computer program that includes computer commands and causes a processor to execute the computer commands, thereby realizing the image orientation identification method described above. [Effects of the Invention]

[0018] The image direction identification method, image direction identification method model, and storage medium provided by the embodiments of the present invention perform fusion on the text line classification results for each text region in the target image to be identified based on the position features and morphological features of the text regions. Compared with the prior art, when the embodiments of the present invention perform fusion on the text line classification results of each text region, they fully consider the position features and morphological features of the text regions, so the noise text in the image is reduced or filtered, thereby improving the accuracy of the image direction identification result. In addition, the embodiments of the present invention use the training data of multiple publicly disclosed regions to pre-train the text region detection module, the feature fusion and text line classification module, reduce the requirements for the model training data, and train the image direction identification module using the training data of the target region, thereby improving the identification performance of the image direction for the target region.

Brief Description of the Drawings

[0019] Hereinafter, other advantages and merits will become clear by explaining the optimal implementation method in detail. The drawings show preferred embodiments, but should not be regarded as limiting the present invention. In the drawings, the same parts are denoted by the same reference numerals. [Figure 1] FIG. 1 is a schematic diagram showing the structure of an image direction identification model according to an embodiment of the present invention. [Figure 2] FIG. 2 is a flowchart showing an image direction identification method according to an embodiment of the present invention. [Figure 3] FIG. 3 is a diagram showing an example of an image direction identification module according to an embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing an example of different text regions according to an embodiment of the present invention. [Figure 5] FIG. 5 is a schematic diagram showing another structure of an image direction identification model according to an embodiment of the present invention. [Figure 6] FIG. 6 is a schematic diagram showing another different structure of an image direction identification model according to an embodiment of the present invention.

Modes for Carrying Out the Invention

[0020] The technical problems, solutions, and advantages of this invention will be described in detail below using drawings and specific embodiments to clarify them further. The provision of specific arrangements and constituent details in the following description is merely to aid in a comprehensive understanding of the embodiments of the invention. Therefore, various changes and modifications can be made to the embodiments described herein, as long as they do not deviate from the scope and spirit of the invention. For the sake of clarity, descriptions of known functions and structures will be omitted.

[0021] Throughout the specification, the terms "one embodiment" or "one example" refer to a specific feature, structure, or characteristic associated with that embodiment, and include it in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in one example" as described in various parts of the specification do not necessarily refer to the same embodiment. These specific features, structures, or characteristics can be combined in any way appropriate in one or more embodiments.

[0022] In each embodiment of the present invention, the magnitude of the process numbers below does not indicate the order of execution. The execution order of each process is determined by its function and inherent logic and is not limited in any way to the implementation processes of the embodiments of the present invention.

[0023] Embodiments of the present invention provide an image orientation identification method for identifying the orientation of an image to be identified. Here, the orientation of an image is typically one of 0°, 90°, 180°, or 270°. For example, 0° indicates that the image is positive, while 90°, 180°, and 270° indicate rotations from 0° to 90°, 180°, and 270°, respectively. In embodiments of the present invention, an image being positive means that the orientation of the characters in the image conforms to the reading habits of the characters. In some cases, there may be multiple character orientations in an image. In this case, the primary character orientation in the image is defined as the character orientation in that image. The primary character is typically the character of the text that occupies the largest area in the image and is oriented in the same direction.

[0024] Similarly, the orientation of text lines in an image is usually one of 0°, 90°, 180°, or 270°. Here, 0° represents the forward direction. 90°, 180°, and 270° represent text lines rotated 90°, 180°, and 270° from 0°, respectively. In embodiments of the present invention, the forward orientation of text lines means that the character direction in the text lines conforms to reading habits.

[0025] In an embodiment of the present invention, an image orientation recognition model is pre-trained. As shown in Figure 1, the image orientation recognition model mainly includes a pre-trained character region detection module 11, a feature fusion and text line classification module 12, and an image orientation recognition module 13. When identifying the orientation of a target image containing characters (the target image to be identified), the target image is input to the image orientation recognition model, and the image orientation recognition model outputs an image orientation identification result, for example, one of 0°, 90°, 180°, or 270°. Alternatively, the image orientation identification result may be a probability value of the above angles.

[0026] Specifically, the character region detection module 11 detects character regions in the target image to be identified and obtains the position coordinates and region features of the character regions. The position coordinates of the character regions are represented by the coordinates of a rectangle that circumscribes the character region, for example, the coordinates of two diagonals of the circumscribed rectangle (for example, the upper left corner and the lower right corner).

[0027] The feature fusion and text line classification module 12 identifies image features in the target image of the character region based on the position coordinates of the character region, fuses them with the region features of the character region to generate a first fused feature; and generates a text line classification result of the character region based on the first fused feature. Here, the method of fusing the image features and region features may be a method such as concatenating or weighted addition of feature vectors, and the embodiments of the present invention are not specifically limited thereto. Furthermore, the text line classification result is one of the directions of the text line, for example, 0°, 90°, 180°, or 270°.

[0028] The image orientation recognition module 13 acquires the positional and morphological features of each character region, and for each character region, it fuses the positional and morphological features with the text line classification results to obtain the orientation recognition result of the target image. The orientation recognition result of the target image may be any of 0°, 90°, 180°, or 270°, or it may be a probability value corresponding to an angle.

[0029] The image orientation identification method includes the following steps, as shown in Figure 2.

[0030] In step 21, the character region is detected from the target image to be identified, and the position coordinates and region features of the character region are obtained.

[0031] Here, the target image to be identified is input to the character region detection module 11, and the position coordinates and region features of each character region output by the character region detection module 11 are obtained. Specifically, the character region detection module generates an image feature diagram corresponding to the target image, detects character regions based on the image feature diagram to identify the position coordinates of the character regions, and further extracts features corresponding to the character regions from the image feature diagram to be used as region features.

[0032] In step 22, based on the position coordinates of the character region, image features in the target image of the character region are identified and fused with the region features to generate a first fused feature; and based on the first fused feature, text line classification results are generated.

[0033] Here, the target image, the position coordinates of the text region, and the region features are input to the feature fusion and text line classification module 12, and the text line classification result of the text region output by the feature fusion and text line classification module 12 is obtained.

[0034] In step 23, positional and morphological features are acquired for each character region, and based on these features, the text line classification results are fused to form the direction identification result for the target image.

[0035] Here, the position coordinates and text line classification results for each character area are input to the image orientation identification module 13, and the image orientation identification module 13 outputs the orientation identification result for the target image. The image orientation identification result is an angle corresponding to the direction, for example, 0°, 90°, 180°, or 270°. Alternatively, the image orientation identification result may be a probability value of the above angles.

[0036] The image orientation recognition module is used specifically in the following ways, namely, For each character area, Based on position coordinates, obtain positional and morphological features of the character region; merge the positional and morphological features to generate a second fused feature; and generate weights for the character region based on the second fused feature; Based on the aforementioned weights, the text line classification results are merged to generate the direction identification result for the target image.

[0037] In embodiments of the present invention, the positional features include a relative polar angle, which is the difference between a polar angle and a preset reference polar angle, and a polar diameter. Specifically, the polar angle and polar diameter are those of predetermined angles (e.g., the upper left corner, the lower right corner) of a rectangle circumscribing the character area, or those of predetermined points (e.g., the center point, the center of mass, or another point) in the character area. In embodiments of the present invention, the polar angle and polar diameter of the character area may be those of the polar angle range and polar diameter range that the character area spans. For example, the minimum or maximum value of the polar angle of the character area, or the minimum and maximum value of the polar diameter of the character area. The morphological features of the character area include at least one of the following: the height, width, area, contour length, or center of mass (e.g., the position coordinates of the center of mass).

[0038] Embodiments of the present invention improve image orientation recognition performance by using an attention mechanism to draw the model's attention to the positional and morphological features of character regions in an image. Figure 3 shows an example of an image orientation recognition module. In this example, a Transformer model is used for the image orientation recognition module. In this model, morphological features include features such as the height, width, and area of ​​the character region. Embodiments of the present invention consider that conventional XY coordinate axes cannot accommodate the need for image features to match as closely as possible even when the image is rotated at various angles, and therefore display the positional features of the character region using position coding based on polar coordinates. The positional characteristics of the character region include the polar diameter and polar angle of the character region. Of these, the polar diameter is encoded using a normalized coding method, and the polar angle can be represented by the difference (relative polar angle) between the polar angle of each character region (token) and the polar angle of the CLS flag (reference polar angle).

[0039] Through the steps described above, the embodiment of the present invention performs fusion on the text line classification results of character regions in a target image, which are identified based on the positional and morphological features of the character regions. Because the positional and morphological features of the character regions are fully considered during the fusion of the text line classification results, noisy text in the image is reduced or filtered, and the accuracy of the image orientation identification results is improved.

[0040] For example, in Figure 4, there are multiple text regions enclosed by red dotted lines. The text lines in these text regions face different directions. For example, the text lines in text regions 41-43 face 0°, while the text lines in text regions 44-45 face 270°. In this embodiment of the present invention, in step 23, a pre-trained image orientation recognition module is used to fuse the text line classification results of each text region based on its positional and morphological features. During training, when learning the weights of the text regions, the image orientation recognition module assigns relatively low weights to text regions at the image edges. Therefore, since the positions of dashed frames 44 and 45 are more biased towards the image edges, the text line classification results within these frames are given relatively low weights during fusion, reducing the influence of noise text in these areas. In contrast, dashed frame 41 has a large area and is therefore assigned a relatively large weight. Also, dashed frame 42 is closer to the center of the image and is therefore assigned a relatively large weight. In this way, by appropriately assigning weights to the text line classification results of the character region according to the positional and morphological features of the character region, the accuracy of the image orientation recognition results is improved.

[0041] The training of the image orientation recognition model according to an embodiment of the present invention will be described below.

[0042] In this embodiment of the present invention, the image orientation recognition model is pre-trained before step 21 described above. Here, a large amount of publicly available data sets already exist for character region detection and text line classification tasks in various fields. Therefore, to reduce the need to represent the data, the character region detection module and the feature fusion and text line classification modules are pre-trained based on existing data sets in many fields. Furthermore, the image orientation recognition module can be improved by performing precise training (fine-tuning) using training data from a specific field. This training method in this embodiment of the present invention reduces the need for training data from a specific field, and by pre-training with many general-purpose data sets, the model can acquire even better prior knowledge. Specifically, the above training includes the following steps.

[0043] (1) The character region detection module is pre-trained using first training data that includes first images of multiple marked character regions (multiple first images in which character regions are marked). Here, the first training data is obtained from publicly available data sets in multiple fields.

[0044] (2) A second set of training data, including a second image of multiple marked text line classification results (a second image in which the text line classification results are marked), is input to a pre-trained character region detection module. The position coordinates and region features of the character regions output by the trained character region detection module, along with the second set of training data, are used to pre-train a feature fusion and text line classification module. Here, the second set of training data may be obtained from publicly available data sets in multiple fields.

[0045] (3) Acquire third training data for the target region. The third training data includes a third image with multiple image orientations marked (multiple third images with image orientations marked). Here, the third training data is usually training data for a specific field. Subsequently, the third training data is input into a pre-trained character region detection module, and the position coordinates and region features of the character region output by the character region detection module, along with the third training data, are input into a pre-trained feature fusion and text line classification module, and the text line classification results output by the feature fusion and text line classification module, along with the third training data, are used to train the image orientation discrimination module.

[0046] As can be seen from the above, the embodiment of the present invention reduces the requirements for model training data by pre-training the character region detection module and the feature fusion and text line classification modules using conventionally published training data from multiple fields. Furthermore, the image orientation recognition module is improved by training it using training data of the target region to recognize the image orientation of the target region.

[0047] Based on the above, the embodiment of the present invention further provides an image orientation recognition model for implementing the above method. Figure 1 shows an example of the configuration of the image orientation recognition model.

[0048] The aforementioned image orientation identification module 13 further, For each character region, positional features and morphological features are identified and merged based on positional coordinates to generate a second merged feature, and weights for the character region are generated based on this second merged feature; By merging the text line classification results based on the weights of the aforementioned character regions, the direction identification result of the target image is obtained.

[0049] Here, the positional features of the character region include the relative polar angle, which is the difference between the polar angle and the reference polar angle, and the polar diameter.

[0050] The morphological features of a character region include at least one of the character region's height, width, area, contour length, and center of mass.

[0051] Furthermore, the character region detection module 11 generates an image feature diagram corresponding to the target image, detects character regions based on the image feature diagram, identifies the position coordinates of the character regions, and extracts features corresponding to the character regions from the image feature diagram to define the region features of the character regions.

[0052] Furthermore, another image orientation recognition model according to an embodiment of the present invention is provided. As shown in Figure 5, this image orientation recognition model further includes a training module 14. The training module 14 is A character region detection module is pre-trained using first training data containing multiple first images marked within character regions (multiple first images in which character regions are marked); The second training data, which includes multiple second images (multiple second images marked with the text line classification result) indicated in the text line classification result, is input to a pre-trained character region detection module, and the feature fusion and text line classification modules are pre-trained using the position coordinates and region features of the character regions output by the character region detection module and the second training data; A third training data set is obtained, which includes multiple third images in the target region with image orientations indicated; The third training data is input to a pre-trained character region detection module, and the position coordinates and region features of the character regions output by the character region detection module, along with the third training data, are input to a pre-trained feature fusion and text line classification module. The image orientation recognition module is then trained using the text line classification results output by the feature fusion and text line classification module, along with the third training data.

[0053] Figure 6 shows the hardware configuration of an image orientation recognition model according to an embodiment of the present invention. As shown in Figure 6, the image orientation recognition model 600 comprises a processor 602 and a memory 604 in which computer program commands are stored, and by causing the processor 602 to execute the computer program commands, The system detects character regions from the target image to be identified, and obtains the position coordinates and region features of each character region; Based on the position coordinates, the image features in the target image of the character region are identified and merged with the region features to generate a first merged feature of the character region; and based on the first merged feature, the text line classification result of the character region is generated; For each character region, positional features and morphological features are acquired, and based on the positional features and morphological features, the text line classification results are merged to generate the direction identification results for the target image.

[0054] Furthermore, as shown in Figure 6, the model training device 600 further includes a network interface 601, an input device 603, a hard disk 605, and a display device 606.

[0055] These interfaces and devices are interconnected via a bus architecture. The bus architecture can be a bus and bridge that can include any number of interconnections. Specifically, it connects various circuits of one or more central processing units (CPUs) and / or graphics processing units (GPUs), represented by processor 602, and one or more memory units, represented by memory 604. The bus architecture can also connect various other circuits, such as peripherals, voltage regulators, and power management circuits. It should be understood that the bus architecture is used to enable communication between these components. In addition to the data bus, the bus architecture includes power buses, control buses, and state signal buses, all of which are publicly known in the art and therefore will not be described in detail here.

[0056] The network interface 601 can connect to a network (e.g., the Internet, a local area network, etc.), receive data such as training data from the network, and save the received data to the hard disk 605.

[0057] The input device 603 can receive various commands input by the operator and send them to the processor 602 for execution. The input device 603 may include a keyboard or a click device (for example, a mouse, trackball, touch panel, or touchscreen).

[0058] The display device 606 can display the results obtained by the processor 602 when it executes commands, for example, the training progress of a model.

[0059] Memory 604 stores programs and data necessary for the operation of the operating system, as well as data such as intermediate results from calculations performed by processor 602.

[0060] In embodiments of the present invention, memory 604 is volatile memory or non-volatile memory, or includes both volatile and non-volatile memory. Among these, non-volatile memory is read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random-access memory (RAM) used as an external cache. Memory 604 of the apparatus and methods described herein includes, but is not limited to, these memories and any other suitable type of memory.

[0061] In some implementations, memory 604 stores an operating system (OS) 6041 and an application program (APP) 6042 as executable modules or data structures, subsets thereof, or extensions thereof.

[0062] Within this system, the operating system 6041 includes various system programs, such as a framework layer, core library layer, and driving layer, and is used to implement various core business operations and hardware-based tasks. The application program 6042 includes various application programs, such as a web browser, and is used to implement various application operations. The program that executes the method according to this embodiment is included in the application program 6042.

[0063] The methods according to embodiments of the present invention are applied to or implemented by a processor 602. The processor 602 is a type of integrated circuit chip having the function of processing signals. In the implementation process, each step of the method is implemented by hardware integrated logic circuits or software-based instructions within the processor 602. The processor 602 is a general-purpose processor, a digital signal processing unit (DSP), an application-designed integrated circuit (ASIC), a commercially available programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, which can implement or execute each method, step, and logic box disclosed in embodiments of the present invention. General-purpose processors include microprocessors or any general-purpose processors. Each step of the methods according to embodiments of the present invention may be implemented by a hardware decoder, or by a combination of hardware and software that can be implemented in the decoder. The software module is stored in a storage medium mature in the art, such as random memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, or registers. The processor 602 reads information from the memory 604, which has a storage medium in which this software is stored, and implements the steps of the above method according to the hardware.

[0064] The embodiments described above are implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. In the case of hardware implementation, the processing unit is implemented using one or more dedicated integrated circuits (ASICs), digital signal processing processors (DSPs), digital signal processing devices (DSPDs), programmable logic circuits (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units that perform the functions of the present invention, or a combination thereof.

[0065] Furthermore, the software implementation is carried out by modules (such as processes and functions) that implement the functions described above. The software code is stored in memory and executed by the processor. Memory can be implemented either internally or externally to the processor.

[0066] Specifically, the computer program is executed on the processor 602, which inputs the position coordinates of each character area and the text line classification results to the image orientation identification module, and further realizes the step of obtaining the orientation identification result of the target image output by the image orientation identification module.

[0067] The aforementioned image orientation identification module is For each character region, the positional features and morphological features are identified based on the positional coordinates, the positional features and morphological features are merged to generate a second fused feature, and the weights of the character region are generated based on the second fused feature; Based on the weight of each character region, the text line classification results for that character region are merged to generate the direction identification result for the target image.

[0068] Furthermore, specifically, when causing the processor 602 to execute the aforementioned computer program command, The aforementioned positional features include the relative polar angle, which is the difference between the polar angle and the reference polar angle, and the polar diameter. The aforementioned morphological features are at least one of the height, width, area, contour length, and center of mass of the character region.

[0069] Furthermore, the process involves causing the processor 602 to execute the computer program command, thereby inputting the target image to the character region detection module, and acquiring the position coordinates and region features of each character region output by the character region detection module.

[0070] Furthermore, the process involves causing the processor 602 to execute the computer program command, thereby inputting the position coordinates and region features of the target image and the character region into a feature fusion and text line classification module, and generating a text line classification result for the character region output by the feature fusion and text line classification module.

[0071] More specifically, by causing the processor 602 to execute the aforementioned computer program command, The character region detection module is pre-trained using first training data, which includes multiple first images on which character regions are marked; Second training data, including multiple second images on which the text line classification results are indicated, is input to the trained character region detection module; the feature fusion and text line classification module is pre-trained using the position coordinates and region features of the character regions output by the character region detection module and the second training data; A third training data set is obtained, which includes multiple third images in the target region with image orientations indicated; The third training data is input to the trained character region detection module, and the position coordinates and region features of the character regions output by the character region detection module, along with the third training data, are input to the trained feature fusion and text line classification module, and the image orientation discrimination module is trained using the text line classification results output by the feature fusion and text line classification module and the third training data.

[0072] An embodiment of the present invention further provides a computer-readable storage medium in which computer program commands are stored, and by causing a processor to execute the computer program commands, the image orientation identification method is realized and the same effect is achieved. A detailed explanation is omitted to avoid redundancy. Among these, the computer-readable storage medium is, for example, ROM (Read-Only Memory), RAM (Random Access Memory), a disk, or an optical disk.

[0073] Embodiments of the present invention further provide a computer program including computer program instructions, and when these computer program instructions are executed by a processor, the image orientation identification method is realized and the same effect can be achieved. To avoid redundancy, a detailed explanation is omitted here.

[0074] In this text, “includes,” “includes in,” or other variations are intended to cover non-exclusive inclusion, so that a process, method, article, or apparatus containing a set of elements includes not only those elements but also other elements not explicitly stated, or elements specific to such a process, method, article, or apparatus. Unless otherwise limited, the elements limited by the word “includes” do not preclude the existence of further identical elements in a process, method, article, or apparatus containing that element.

[0075] From the embodiments described above, it is clear that the methods of the embodiments are implemented via software and the necessary general-purpose hardware platform. Of course, they can also be implemented in hardware, but in many cases the former is a superior embodiment. With this understanding, the parts of the present invention that are essential or contribute to the prior art are expressed in the form of a software product. The computer program product is stored on a storage medium (e.g., ROM / RAM, disk, optical disk) and includes computer program instructions that cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method of each embodiment of the present invention.

[0076] Although embodiments of the present invention have been described above with reference to the drawings, the present invention is not limited to specific embodiments. The specific embodiments described above are merely illustrative and not limiting. Modifications and changes can be made as long as they do not depart from the spirit and claims of the present invention, but all such modifications will remain within the scope of the present invention.

Claims

1. A computer-based method for determining the orientation of an image, Multiple character regions are detected from the target image to be identified, and the position coordinates and region features of each character region are obtained; For each character region, the image features in the target image of the character region are identified based on the position coordinates, the image features and the region features are merged to generate a first merged feature of the character region, and based on the first merged feature, a plurality of text line classification results for the character region are generated; This includes acquiring positional and morphological features for each character region, fusing the classification results of multiple text lines in the character region based on the positional and morphological features, and generating a direction identification result for the target image. A method for identifying the orientation of an image, characterized by the features described above.

2. For each of the aforementioned character regions, positional features and morphological features are acquired, and based on the positional features and morphological features, the classification results of multiple text lines in the character region are merged to generate the direction identification result of the target image. The process includes inputting the position coordinates of each character area and the text line classification result into an image orientation identification module, and obtaining the orientation identification result of the target image output by the image orientation identification module, The aforementioned image orientation identification module is For each character region, the positional features and morphological features are identified based on the positional coordinates, the positional features and morphological features are merged to generate a second fused feature, and the weights of the character region are generated based on the second fused feature; Based on the weight of each character region, the classification results of multiple text lines in that character region are merged to generate the direction identification result of the target image. The image orientation identification method according to feature 1.

3. The aforementioned positional features include the relative polar angle, which is the difference between the polar angle and the reference polar angle, and the polar diameter. The aforementioned morphological features include at least one of the height, width, area, contour length, and center of mass of the character region. The image orientation identification method according to feature 2.

4. Detecting multiple character regions from the target image and obtaining position coordinates and region features for each character region is possible. This includes inputting the target image into a character region detection module and obtaining the position coordinates and region features of each character region output by the character region detection module. The image orientation identification method according to feature 2.

5. For each character region, the image features of the target image of the character region are identified based on the position coordinates, the image features and the region features are merged to generate a first merged feature of the character region, and a plurality of text line classification results of the character region are generated based on the first merged feature. This includes inputting the position coordinates and regional features of the target image and the character region into a feature fusion and text line classification module, and generating multiple text line classification results for the character region output by the feature fusion and text line classification module. The image orientation identification method described in feature 4.

6. Before detecting the character region from the aforementioned target image, The character region detection module is pre-trained using first training data that includes multiple first images on which character regions are marked; Second training data, including multiple second images on which the text line classification results are indicated, is input to the trained character region detection module; the feature fusion and text line classification module is pre-trained using the position coordinates and region features of the character regions output by the character region detection module, and the second training data; A third training data set is obtained, which includes multiple third images in the target region with their orientations indicated; The third training data is input to the trained character region detection module, and the position coordinates and region features of the character regions output by the character region detection module, along with the third training data, are input to a trained feature fusion and text line classification module, and the image orientation recognition module is trained using the text line classification results output by the feature fusion and text line classification module and the third training data. The image orientation identification method according to feature 5.

7. A character region detection module that detects multiple character regions from an image to be identified and obtains the position coordinates and region features of each character region, A feature fusion and text line classification module that, for each character region, identifies image features in the target image of the character region based on the position coordinates, fuses the image features and the region features to generate a first fused feature of the character region, and generates multiple text line classification results for the character region based on the first fused feature, Includes an image orientation recognition module that acquires positional features and morphological features for each of the character regions, merges the classification results of multiple text lines in the character region based on the positional features and morphological features, and generates an orientation recognition result for the target image. An image orientation identification device characterized by the following:

8. The aforementioned image orientation identification module further, For each character region, the positional features and morphological features are identified based on the positional coordinates, the positional features and morphological features of the character region are merged to generate a second fused feature, and the weight of the character region is generated based on the second fused feature; Based on the weight of each character region, the classification results of multiple text lines in that character region are merged to generate the direction identification result of the target image. The image orientation identification device according to feature 7.

9. The aforementioned positional features include the relative polar angle, which is the difference between the polar angle and the reference polar angle, and the polar diameter. The aforementioned morphological features include at least one of the height, width, area, contour length, and center of mass of the character region. The image orientation identification device according to feature 8.

10. The aforementioned character area detection module further, An image feature diagram corresponding to the target image is generated, the character region is detected based on the image feature diagram, the position coordinates of the character region are obtained, and features corresponding to the character region are extracted from the image feature diagram and used as the region features of the character region. The image orientation identification device according to feature 8.

11. The character region detection module is pre-trained using first training data that includes multiple first images on which character regions are marked; Second training data, including multiple second images on which the text line classification results are indicated, is input to the trained character region detection module; the feature fusion and text line classification module is pre-trained using the position coordinates and region features of the character regions output by the character region detection module, and the second training data; A third training data set is obtained, which includes multiple third images in the target region with their orientations indicated; The present invention further includes a training module which inputs the third training data into the trained character region detection module, inputs the position coordinates and region features of the character regions output by the character region detection module, and the third training data into a trained feature fusion and text line classification module, and trains the image orientation discrimination module using the text line classification results output by the feature fusion and text line classification module and the third training data. The image orientation identification device according to feature 10.

12. A program for causing a computer to perform the image orientation identification method described in any one of claims 1 to 6.

13. A computer-readable storage medium storing the program described in claim 12.