Method for analyzing at least one image, electronic device and corresponding computer program product

The method improves text analysis by extending image blocks based on spacing criteria to merge adjacent graphic objects, addressing the inefficiencies of existing algorithms and reducing computational costs, thereby enhancing text body recognition.

FR3155939A1Pending Publication Date: 2025-05-30ORANGE SA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
FR2023013091
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing automatic text detection algorithms struggle to efficiently group words into more global text bodies, such as lines or paragraphs, due to their independent nature and high computational costs, often requiring large training corpora for effective neural network-based methods.

Method used

A method for analyzing images by delimiting blocks of pixels encompassing graphic objects likely to represent alphanumeric characters, and then extending these blocks in specific directions to merge adjacent graphic objects based on spacing criteria, facilitating temporary storage for character recognition.

Benefits of technology

This approach allows for more efficient grouping of image portions into text bodies, reducing computational costs and eliminating the need for large training datasets, thereby enhancing the accuracy and efficiency of text analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for analyzing at least one image, electronic device and corresponding computer program product The present application relates to a method for analyzing an image, implemented by an electronic device and comprising: a delimitation, in an image, of blocks of pixels encompassing graphic objects capable of corresponding to alphanumeric characters; at least one extension of a first of the blocks in a first direction of extension to encompass, in the first block, both a first graphic object already encompassed in the first block and at least one second graphic object of at least one second of the blocks of pixels, the extension in the first direction of extension being implemented when the first block of pixels is spaced in the first direction of extension by less than a first number of pixels from a second block of pixels, the first number of pixels taking into account a height of at least one first graphic object of the first block of pixels;storage of the first extended block for character recognition. The present application also relates to an electronic device implementing such a method as well as the corresponding computer program and recording medium. Figure for the abstract: Fig. 3;
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for analyzing at least one image, electronic device and corresponding computer program product 1. Technical field

[0001] The present application relates to the field of automatic image analysis to identify elements likely to represent typographic (alphanumeric) characters. It relates in particular to a method for at least partially automatic analysis of at least one image, implemented by at least one electronic device, as well as a corresponding electronic device, computer program product and medium. 2. State of the art

[0002] Automatic text detection aims to locate in an image portions of the image likely to contain textual elements. It generally precedes other tasks such as the recognition of characters, or words, actually present in these portions of the image, and the analysis of texts formed by these words (for example a semantic analysis).

[0003] Some text detection algorithms that seek to locate words in an image begin by identifying in this image portions of the image that are likely to correspond to words, based on the difference, in terms of length, between inter-word spaces and inter-character spaces. Indeed, two consecutive characters belonging to different words are generally further apart than two consecutive characters of the same word. This approach thus identifies portions of the image that are independent of each other, each corresponding to a word. As a result, the words extracted by such algorithms are independent of each other, which prevents other words already detected in other portions of the image from being taken into account when interpreting the current portion of the image.

[0004] Thus, it is necessary to add another step, after the recognition of the words of the image portions, in order to group the words into more global text bodies (lines or paragraphs).

[0005] Methods have been developed to group words together. Some methods are based, for example, on the use of neural networks. However, such methods are very expensive in terms of computational time and, to be effective, these methods often require very large training corpora for training these neural networks.

[0006] The object of the present application is to propose improvements to at least certain disadvantages of the state of the art. 3. Statement of the invention

[0007] The present application aims to improve the situation using a method for analyzing at least one image, implemented by at least one electronic device, said method comprising: - a delimitation, in an image, of blocks of pixels encompassing graphic objects of the image likely to correspond to alphanumeric characters; - at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least a first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block for the purpose of implementing character recognition on said first extended block.

[0008] According to at least one embodiment, said extension of said first block in said first direction of extension includes several second blocks located consecutively in said first direction of extension, two second consecutive blocks being spaced from each other by less than a second number of pixels, said second number of pixels taking into account a height of the second consecutive block closest to said first block in said first direction of extension.

[0009] According to at least one embodiment, said extension of said first block in said first direction of extension encompasses all the blocks located in said first direction of extension, starting from said first block, and spaced from each other by less than said first number of pixels.

[0010] According to at least one embodiment, the method comprises at least one extension of said first block in a second extension direction to encompass in said first block both said first and second graphic objects already encompassed in said first block and at least one third graphic object of at least one third block of pixels, said at least one extension in said second extension direction being implemented when said first block of pixels is spaced in a second extension direction by less than a third number of pixels from one of said at least one third block of pixels, said third number of pixels taking into account a height of at least one graphic object of said at least one first block of pixels

[0011] According to at least one embodiment, said extension of said first block in said second direction of extension includes several third blocks located consecutively in said second direction of extension, two consecutive third blocks being spaced from each other by less than a third number of pixels, said third number of pixels taking into account a height of the third consecutive block closest to said first block in said first direction of extension.

[0012] According to at least one embodiment, said extension of said first block in said second direction of extension encompasses all the blocks located in said second direction of extension, starting from said first block, and spaced from each other by less than said third number of pixels.

[0013] According to at least one embodiment, said extension according to said second direction of extension is implemented conditionally, when said first block comprises all the graphic objects detected in said image aligned with said first graphic object according to said first direction of extension.

[0014] According to at least one embodiment, said method comprises an association with said first graphic object of a first character size.

[0015] According to at least one embodiment, said first number of pixels corresponds to an inter-word spacing of a first character font for said first character size.

[0016] According to at least one embodiment, said method comprises an association of said second graphic object with said second consecutive block closest to a second character size and where said second number of pixels corresponds to an inter-word spacing of a second character font for said second character size.

[0017] According to at least one embodiment, said third number of pixels corresponds to an inter-line spacing of said first character font for said first character size.

[0018] According to at least one embodiment, the method comprises obtaining a designation of said first direction of extension.

[0019] The characteristics presented in isolation in the present application in connection with certain embodiments of the method of the present application can be combined with each other according to other embodiments of the present method.

[0020] According to another aspect, the present application also relates to an electronic device adapted to implement the method of the present application in any of its embodiments. For example, the present application thus relates to an electronic device comprising at least one processor configured for an analysis of at least one image, said analysis comprising: - a delimitation, in an image, of blocks of pixels encompassing objects image graphics that may correspond to alphanumeric characters; - at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least a first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block for the purpose of implementing character recognition on said first extended block.

[0021] The present application also relates to a computer program comprising instructions for implementing the various embodiments of the above method, when the program is executed by a processor, and a recording medium readable by an electronic device and on which the computer program is recorded.

[0022] For example, the present application thus relates to a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for analyzing at least one image, implemented by at least the electronic device, said method comprising: - a delimitation, in an image, of blocks of pixels encompassing graphic objects of the image likely to correspond to alphanumeric characters; - at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least a first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block with a view to updating implementing character recognition on said first extended block.

[0023] For example, the present application also relates to a recording medium readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for analyzing at least one image, implemented by at least the electronic device, said method comprising: - a delimitation, in an image, of blocks of pixels encompassing graphic objects of the image likely to correspond to alphanumeric characters; - at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least a first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block for the purpose of implementing character recognition on said first extended block.

[0024] The programs mentioned above may use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0025] The recording (or information) media mentioned in the present application may be any entity or device capable of storing the program. For example, a medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or even a magnetic recording means.

[0026] Such a storage means may for example be a hard disk, a flash memory, etc.

[0027] On the other hand, an information medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. A program according to the invention may in particular be downloaded from a network such as the Internet.

[0028] Alternatively, an information (or recording) medium may be a circuit integrated circuit in which a program is incorporated, the circuit being adapted to execute or to be used in the execution of any of the embodiments of the method which is the subject of the present patent application.

[0029] Generally speaking, by obtaining an element, is meant in the present application for example a reception of this element from a communication network, an acquisition of this element (via for example user interface elements or sensors), a creation of this element by various processing means such as by copying, encoding, decoding, transformation etc. and / or an access of this element from a local or remote storage medium accessible to at least one device implementing, at least partially, this obtaining. 4. Brief description of the drawings

[0030] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:

[0031] [Fig.l] shows a simplified view of a system, cited as an example, in which at least certain embodiments of the method of the present application can be implemented,

[0032] [Fig.2] shows a simplified view of a device suitable for implementing at least certain embodiments of the method of the present application,

[0033] [Fig. 3] presents an overview of the method of processing the present application, in certain of its embodiments.

[0034] [Fig.4] shows an example of an image to be analyzed by the processing method of the present application, in some of its embodiments;

[0035] [Fig.5] shows as an example the coordinates of a bounding box in a reference frame linked to an image to be analyzed;

[0036] [Fig.6] shows an example of a bounding box as illustrated in [Fig.5] in certain embodiments;

[0037] [Fig.7] shows another example of a bounding box in other embodiments;

[0038] [Fig.8] shows a pixel matrix used for calculating a character size in some embodiments. 5. Description of the embodiments

[0039] The present application proposes a method for grouping (i.e. aggregating) image portions (i.e. blocks of pixels) capable of containing at least one alphanumeric character, to form more global image portions, capable of corresponding to text bodies (such as sets of words, lines or paragraphs). Such a method can thus allow, for example, the subsequent processing of these aggregated image portions taking into account, for the identification of the words of the more global image portions, a context associated with these aggregated image portions, potentially richer than that of the image portions before their aggregation. For example, the method of the present application can thus facilitate the identification of named entities composed of several words, the automatic classification of paragraphs or the application of an automatic natural language processing algorithm.

[0040] The image to which the image portions belong may correspond, for example, to a digitized version of a paper document (scanned for example), or to the content of at least one window displayed on a screen coupled to a device on which the method of the present application is executed, at least partially, or even to a photo or video containing text.

[0041] The present application will now be described in more detail in connection with [Fig. 1].

[0042] [Fig.l] represents a telecommunications system 100 in which certain embodiments of the invention can be implemented. The system 100 comprises one or more electronic devices, at least some of which can communicate with each other via one or more communication networks, possibly interconnected, such as a local area network or LAN (Local Area Network) and / or a wide area network, or WAN (Wide Area Network). For example, the network may comprise a corporate or home LAN network and / or a WAN network of the internet type, or cellular, GSM - Global System for Mobile Communications, UMTS - Universal Mobile Telecommunications System, Wifi - Wireless, etc.).

[0043] As illustrated in [Fig.l], the system 100 may also comprise several electronic devices, such as a terminal (such as a laptop 110, a smartphone 120, a tablet 130), and / or a server 140, for example an application server, a storage device 150, so-called peripheral devices (such as a scanner 160). The system may also comprise management and / or network interconnection elements (not shown).

[0044] [Fig. 2] illustrates a simplified structure of an electronic device 200 of the system 100, for example the device 100, 120, 130 of [Fig. 1], adapted to implement the principles of the present application. Depending on the embodiments, it may be a server, and / or a terminal.

[0045] The device 200 comprises in particular at least one memory M 210. The device 200 may in particular comprise a buffer memory, a volatile memory, for example of the RAM type (for “Random Access Memory” according to English terminology), and / or a non-volatile memory (for example of the ROM type (for “Read Only Memory” according to English terminology). The device 200 may also comprise a processing unit UT 220, equipped for example with at least one processor P 222, and driven by a computer program PG 212 stored in memory M 210. Upon initialization, the code instructions of the computer program PG are for example loaded into a RAM memory before being executed by the processor P. Said at least one processor P 222 of the processing unit UT 220 can in particular implement, individually or collectively, any one of the embodiments of the method of the present application (described in particular in relation to [Fig.3]), according to the instructions of the computer program PG.

[0046] The device may also comprise, or be coupled to, at least one input / output module LO 230, such as a communication module, allowing for example the device 200 to communicate with other devices of the system 100, via wired or wireless communication interfaces, and / or such as a module for interfacing with a user of the device (also called more simply in this application “user interface” or “man-machine interface”).

[0047] By user interface (or “human-machine interface”) of the device, we mean for example an interface integrated into the device 200, or a part of a third-party device coupled to this device by wired or wireless communication means. For example, it may be a secondary screen of the device, or an augmented reality headset connected to the device.

[0048] A user interface may in particular be a user interface, called an “output” interface, adapted to a rendering (or to the control of a rendering) of an output element of a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example on the server 140 of the system 100, or an application accessible via the device 200. Examples of output user interfaces of the device include one or more screens, in particular at least one graphic screen (touch screen for example), a connected headset.

[0049] By rendering, we mean here a restitution (or “output” according to English terminology) on at least one user interface, in any form, for example comprising textual, audio and / or video components, or a combination of such components.

[0050] Furthermore, a user interface may be a so-called “input” user interface, adapted to acquiring a command from a user of the device 200. This may in particular be an action to be performed in connection with a returned item, and / or a command to be transmitted to a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example on the server 140 of the system 100. Examples of input user interfaces of the device 200 include a sensor, an audio acquisition means and / or video (camera (webcam) for example), a keyboard, a mouse.

[0051] Said at least one microprocessor of the device 200 may in particular be adapted for an analysis of at least one image, said analysis comprising: - a delimitation, in an image, of blocks of pixels encompassing graphic objects of the image likely to correspond to alphanumeric characters; - at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least a first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block for the purpose of implementing character recognition on said first extended block.

[0052] Some of the above input-output modules are optional and may therefore be absent from the device 200 in certain embodiments. In particular, if the present application is sometimes detailed in connection with a device communicating with at least one second device of the system 100, the method may also be implemented locally by a device (for example one of the devices of the system 100), for example to analyze an image of a file or document stored locally on the device.

[0053] On the contrary, in some of its embodiments, the method can be implemented in a distributed manner between at least two devices 110, 120, 130, 140 and / or 150 of the system 100.

[0054] The term "module" or the term "component" or "element" of the device is understood here to mean a hardware element, in particular wired, or a software element, or a combination of at least one hardware element and at least one software element. The method according to the invention can therefore be implemented in various ways, in particular in wired form and / or in software form.

[0055] [Fig. 3] illustrates certain embodiments of the method 300 of the present application. The method 300 may for example be implemented by the electronic device 200 illustrated in [Fig. 2].

[0056] As illustrated in [Fig.3], the method 300 may comprise obtaining 310 at least one image to be analyzed (such an image 410 is illustrated in [Fig.4]).

[0057] Depending on the embodiments, this may be at least one image received via a communication interface of the device 200 (from a remote device such as a server, another terminal or even a peripheral such as a scanner or a camera), or an image loaded into memory by the device 200 from a storage medium (database, files, USB key, etc.) accessible to the device 200, or even an image from a “screen copy” of the device 200, the image being able to be read directly in the memory used to manage the display of the device 200.

[0058] As illustrated in Figures 3 and 4, the method 300 may comprise obtaining 320 (or delimiting) blocks 420 of pixels each in association with a graphic object (or content element) detected 321 in the image 410 obtained 310. These content elements may potentially represent one or more alphanumeric characters (such as one or more words, one or more lines of text, etc.). They may also correspond to illustrations other than alphanumeric characters (character recognition not having yet been carried out at this stage).

[0059] In certain embodiments, the detection 321 of at least one content element in the image to be analyzed may comprise the application 3211 of a color transformation to the pixels of the image, so as to obtain pixels in “gray levels” (and therefore a grayed-out image), a filtering 3212 of the transformed image (for example by using a Gaussian blur) to remove certain possible visual noises, that is to say modifying the gray level of pixels, initially different from that of the pixels which are adjacent to them, to give it a value closer to the gray levels of the adjacent pixels (for example a gray level identical or almost identical to the gray levels of these adjacent pixels)

[0060] This filtering 3212 may be optional in certain embodiments.

[0061] The detection 321 of at least one content element may comprise a setting in implementation 3213 of an edge detection algorithm on the grayed and optionally filtered image. The edge detection algorithm used may vary depending on the embodiments. In embodiments where the method of the present application is implemented to isolate content elements likely to be characters (or character strings), the edge detection algorithm used may be an algorithm that has proven reliable on character-type content elements, such as for example the algorithm known as the de Canny algorithm.

[0062] As illustrated in [Fig.3], the method may also comprise an association 322 of at least one block of pixels with at least one detected content element 321. This involves associating with a content element detected in the image a block of pixels encompassing this content element. This block will also be called a “bounding box”. (or “bounding box” in English terminology) in this application, in relation to the content element that it encompasses.

[0063] Obtaining a block of pixels associated with a content element may comprise a fine clipping of the content element (for example as illustrated by an application of a morphological dilation 3221 on a portion, containing the detected content element, of the transformed and filtered image then a new detection 3223 of contours on the dilated image portion. Similar to what was explained previously, the detection 3223 of contours on the dilated image portion may be optionally preceded by a filtering 3222 of visual noise) and implementing a detection algorithm such as Canny.

[0064] Obtaining a block of pixels may comprise a calculation 3224 of positioning information (position, coordinates, orientation, etc.) relating to at least one block of pixels completely encompassing the contours of the clipped content element and a selection of a block (in embodiments where positioning information relating to several encompassing blocks is obtained), for example the block of the smallest size. The calculation 3224 (after possible selection) may for example result in a delimitation of a bounding box of minimal surface area (subject to compliance with certain possible constraints) for this content element. In certain embodiments, the calculation 3224 may take into account constraints (accessible for example via a configuration file or a user interface of the device) limiting certain calculation parameters.Thus, in some embodiments, a constraint may force the orientation of a box, for example to obtain an orientation of the box parallel to an edge of the image (for example horizontal) of a box, regardless of the orientation of the enclosed content element (as illustrated in [Fig.6]). The bounding box will therefore be the minimal area bounding box for this orientation. In other embodiments, no constraint may be relative to the orientation of the box, the orientation of a box then being that allowing the area of ​​the box to be minimized (by aligning for example with the orientation of the content element as illustrated in [Fig.7]).

[0065] This obtaining of block(s) can be carried out for several detected content elements (for example, according to the embodiments, for the content elements located in a portion to be processed of an image or for all the detected content elements).

[0066] The shape of the blocks may vary depending on the embodiments. For example, they may each be rectangular, or each hexagonal, or octagonal, or pentagonal, or ellipsoidal. In some embodiments, they may have different shapes from each other, for example due to the different shapes of the content elements they encompass.

[0067] In embodiments such as the example illustrated in Figure 4, where, for simplicity, a single image 410 is analyzed and where the blocks 420 of pixels are rectangular blocks, each block can for example be represented by 5 positioning variables: .^, y, w, h, 0 where (for this block): - x and y are the coordinates of a point, used as the reference location of the block relative to an origin point of a reference frame.

[0068] For example, in the example of [Fig.4], it may be the coordinates of the upper left corner of the block relative to an origin point corresponding to the upper left corner of the analyzed image. - w is the width of the block - h is the height of the block - 9 is the angle of the block relative to a direction used as a reference (for example p the abscissa axis of the reference frame).

[0069] An analyzed image, in which n blocks B of pixels have been delimited, each represented by 5 positioning variables, can therefore be associated with a set of blocks B of n 5-tuples (with n natural integer greater than or equal to 1) where: B-[B}...Bn], where B^^Xp y., Wp hp with for example 0^]-^, f[vi

[0070] Alternatively, when the analyzed image is a 3-dimensional image, the coordinates of the reference location point of the block will be 3-dimensional coordinates X, y, and Z, each block then being represented by 6 positioning variables.

[0071] Alternatively, when several images are analyzed, the variables used to represent a block may include, in addition to those introduced above, an identifier of the image where the reference location point is located (for example a chronological number, identifying the order of an image in a sequence of images).

[0072] For reasons of simplicity, in order to facilitate the reader's understanding, the remainder of this application will be based on the case of a 2-dimensional reference system having as its point of origin the upper left corner of the analyzed image and having as its axis on the one hand a horizontal abscissa axis and on the other hand a vertical ordinate axis. We also assume an angle 6 — 0 for all the blocks, relative to the horizontal direction (reference direction). In certain embodiments, the method 300 may comprise (as illustrated in [Fig.3]) an association 340 with a block of a character size.

[0073] As highlighted above, the detected content element may not be textual. However, the method 300 of the present application aims to group content elements for subsequent analysis by a character recognition algorithm (for example, an optical character recognition (or OCR) algorithm). for Optical Character Recognition in English). Therefore, in the method 300, the assumption is made of textual content elements (any non-textual content elements can be filtered later during the implementation of this character recognition algorithm).

[0074] The way of obtaining (i.e. estimating) the character size of a content element (assumed to be textual) may vary according to the embodiments.

[0075] Thus, the character size associated with a bounding box can in particular take into account the height of the bounding box.

[0076] For example, in certain embodiments, for a box B^ we can estimate the size size,- of a textual content element potentially present in this box by the height hj of the box.

[0077] As illustrated in FIG. 8, in certain embodiments, (in particular when the boxes are rectangular blocks of pixels) the estimation of the character size may comprise obtaining a matrix Mi 810, relating to a box B, 820, the elements 830 of which correspond to the values ​​(in terms of gray level) of the pixels of the box. An element of the matrix may for example take an integer value between 0 and 255, this value V corresponding to the gray level of the pixel corresponding to this element. For example, the value “V=0” may correspond to “white” and the value “V=255” to black. Depending on the embodiments, it may be a matrix of Aligns and w / columns or, alternatively, a matrix of ^rows and ^columns.

[0078] Optionally, the method may comprise an application of a Gaussian blur on the matrix obtained to eliminate visual noise in the matrix (i.e. the standardization of the value of an isolated pixel with the values ​​of its neighboring pixels). We therefore obtain sequences of pixels of similar colors.

[0079] The estimation of the character size may comprise a sum S (reference 840) of the values ​​V (in gray level) of the pixels 830 of each row (as illustrated in FIG. 8) (respectively each column) of the matrix to obtain a vector V of dimension h, (respectively w;).

[0080] The estimation can also include the calculation of a vector V' of dimension (h--1) representing the variations between two consecutive elements of V (therefore between two columns, respectively two rows, of the matrix M).

[0081] We can express y' by V] = Vj+i - Vj V je[jo, h^}

[0082] The character size can be estimated by the value ( a - b ) where a and b are the indices (with a > b) of the two maximum values ​​of y'.

[0083] Such a calculation may, at least in some embodiments, allow for obtaining an estimate of the character size with better accuracy than simply based on the height of the bounding box.

[0084] Furthermore, such a calculation offers the advantage of being adapted to any color of text and background of the image.

[0085] As illustrated in [Fig.3], the method 300 may comprise obtaining 350 (in number of pixels) of inter-word space sizes and / or obtaining 352 (in number of pixels) of inter-line space sizes taking into account the character size obtained.

[0086] Indeed, the readability of a document requires in practice a consistency between the size of the characters and the inter-word and inter-line spacing of these characters. In particular, the generation of typewritten characters, or those generated via word processing software, obeys typographic standards. Such standards often impose a dependency (for example a proportionality) between a size of character(s) (for example a size of a particular character used as a reference) and the size of the inter-word and inter-line spaces for these characters. For example, according to certain fonts, the size of a space between two words must be less than half the size of the character ''N”.The size occupied by a character (in pixels for example) depends on the font used for this character, obtaining 350, 352 the inter-word and / or inter-line spaces can therefore also take into account, not only the estimated size 320 of the content element but also a font that is assumed to be used by the content element (i.e. the font used as a reference). For example, in certain embodiments, the reference font can be a font from the Office © or Open Office © suite of the Windows © operating system.

[0087] In certain embodiments, as illustrated in [Fig.3], the method 300 may comprise an obtaining 330 of a font assumed to be used (i.e. the reference font on which the calculation will be based). This obtaining 330 may for example correspond to a reading of information designating a font in a configuration file or a reception of such signature information via a user or communication interface of the device 200.

[0088] Once the reference font has been obtained, the value (size) in pixels of the inter-word spaces (denoted AP for block i) and / or the value in pixels (denoted A * for block i) of the inter-line spaces can be obtained 350, 352 by reading a correspondence table, associating reference fonts and character sizes (or ranges of character sizes) with a respective value of inter-word or inter-line space. It should be noted that these are values ​​considered as maximum spacing between words or between lines of the same paragraph.

[0089] In some embodiments, the font corresponding to the largest inter-word or inter-line spaces, in terms of number of pixels, may be used by default as the reference font, so as to limit the risk of not gathering in the same block two characters of the same word, respectively of the same paragraph.

[0090] For example, the method may comprise a delimitation 360, 362 of at least one temporary block each corresponding to an expansion of an obtained block 320, according to at least one direction of extension, as a function of the inter-word and / or inter-line spaces associated with this obtained block 320.

[0091] As illustrated in [Fig.3], the method may comprise a delimitation 360, 362 of temporary blocks for all the obtained blocks 320 (for example a joint delimitation of temporary blocks for all the obtained blocks).

[0092] Thus, for a block obtained 320 (also called hereinafter block considered), the method may comprise in certain embodiments a delimitation 360 of a temporary block, corresponding to an expansion in width of the block considered, of a width equal to the inter-word space A ” obtained 350, the delimitation 360 leading to a temporary block of height and angle equal to those of the block considered, and of width w'= w+ A f with the notations used above (or alternatively of width slightly greater than this width w').

[0093] In certain embodiments, the method may comprise a delimitation 362 of a temporary block, corresponding to an expansion in height of the block considered, of a height equal to the calculated inter-line space A, the delimitation 362 leading to a temporary block of width and angle equal to those of the current block, and of height h' = h+ A • with the notations used above (or alternatively of height slightly greater than this height h').

[0094] In certain embodiments, the delimitation of a temporary block may be obtained by an expansion in two directions at the same time of the corresponding enclosing block, thus resulting in a temporary block of the same angle, of width w'= w+ A and of height h' = h+ A with the notations above.

[0095] In some embodiments, the expansion along an extension direction can be done in the same orientation for all the enclosing blocks. Thus, according to a first example, all the temporary blocks can be obtained by a horizontal expansion to the right (respectively to the left) of the enclosing boxes. According to a second example, all the temporary blocks can be obtained by a vertical expansion downwards (respectively upwards) of the enclosing boxes. According to a third example, all the temporary blocks can be obtained by a horizontal expansion to the right (respectively to the left) of the enclosing boxes and a vertical expansion downwards (respectively upwards).

[0096] The method may also comprise a verification 370, 372 of an overlap between the delimited temporary blocks and at least one enclosing block (neighboring these temporary blocks). It is noted that in embodiments where the expansion is carried out (for all temporary blocks) in at least one direction according to a single orientation (only to the left or only to the right for example for a horizontal direction), this amounts to looking at an overlap between a temporary block and the neighboring temporary block according to this direction and this orientation.When an overlap exists (370) between a temporary block, delimited by width expansion (360) with respect to the corresponding enclosing block, and another enclosing block, neighboring this enclosing block corresponding to the temporary block according to this direction and this orientation, this means that the two enclosing blocks are separated by a distance less than the inter-word space A f and therefore that the content elements of the two blocks are separated by a distance less than the maximum inter-word space within the same paragraph (or more generally, to generalize the explanation to enclosing blocks potentially containing several content elements, that the content elements which are closest (i.e. the rightmost content element of one of the blocks and the leftmost content element of the other block or vice versa) are separated by a distance less than the maximum inter-word space within the same paragraph).Therefore, if these content elements represent characters, they logically belong to the same word or to different words in the same paragraph.

[0097] Similarly, when an overlap 372 exists between a temporary block, delimited by height expansion (362) relative to the corresponding enclosing block, and another enclosing block, neighboring this enclosing block corresponding to the temporary block according to this direction and this orientation, this means that the two enclosing blocks are separated by a distance less than the inter-line space a * and therefore that the content elements of the two enclosing blocks are separated by a distance less than the maximum inter-line space within the same paragraph (or more generally, to generalize the explanation to blocks potentially containing several content elements,that the content elements of the two enclosing blocks that are closest (i.e. the topmost content element of one of the blocks and the bottommost content element of the other block or vice versa) are separated by a distance less than the maximum interline space within the same paragraph). Therefore, if these content elements represent characters, they logically belong to the same paragraph. As illustrated, when an overlap exists for at least one temporary block, the method 300 comprises an extension 380 of the enclosing block corresponding to this temporary block to encompass at least one other enclosing block that overlaps this at least one temporary block.

[0098] The steps of delimitation 360, 362, verification 370, 372 of overlap and extension 380, 382 may differ according to the embodiments. For example, in certain embodiments, a delimitation 360, 362 can be carried out according to two directions of extension on the set of enclosing blocks obtained, an overlap verification then an extension being then carried out globally by set of blocks overlapping in these two directions of extension, or alternatively by set of blocks overlapping in one direction of extension, then optionally (or conditionally) in a second direction of extension. In other embodiments, a delimitation 360, 362 can be carried out according to two directions of extension on the set of enclosing blocks obtained, an overlap verification then an extension being then carried out iteratively, block by block (in reiteration on a block already extended), in two directions of extension, or alternatively in the same direction of extension, then optionally (or conditionally) in the same second direction of extension.It can also involve successive or joint delimitation / verification / extension in different extension directions and / or orientations (for example left then right, right then up, etc.).

[0099] A first example is presented below where a delimitation 360, 362 is made on the set of enclosing blocks in two directions of extension, a first overlap verification 370 and a first extension 380 being made “globally”, by set of enclosing blocks for which there is an overlap (via their temporary blocks) in a first direction of extension (in width (horizontally) towards the right in our example), this first extension being followed by a second overlap verification 372 and a second “global” extension 382 also, by set of enclosing blocks (therefore resulting from the first extension) for which there is an overlap (via their temporary blocks) in a second direction of extension (vertically downwards in our example).In such an embodiment, when an overlap exists 370, 372 for a temporary block, the method 300 comprises an extension 380, 382 of the enclosing block corresponding to this temporary block to encompass both its own content element(s) (or alternatively the content elements of this temporary block).

[0100] and the content element(s) of all enclosing blocks encountered, starting from it in the chosen extension direction and orientation, and overlapped by a temporary block (or alternatively the content elements of these temporary blocks).

[0101] If the temporary blocks are delimited by expansion in a single orientation per direction of extension, and if we denote by E the set of enclosing blocks, including the enclosing block Bz, whose temporary blocks overlap two by two (i.e. have a non-zero intersection), the enclosing block Bz once extended therefore corresponds to a fusion of the blocks of this set Ei.

[0102] We then obtain, for verifications and extensions 380 in width, for each set of blocks overlapping in width, a block of approximate width the sum of the widths of these merged blocks, to which is added the inter-word distance(s) between these blocks, and the height of which is the greatest of the respective heights of the blocks.

[0103] Similarly, after verification 372 of overlap in height and an extension 382 in height, we obtain for each set of blocks overlapping in height, a block of approximate height the sum of the heights of the merged blocks, to which is added the inter-line distance(s) between these blocks, and the width of which is the largest of the respective widths of these blocks.

[0104] A second example is now described where first delimitations 360, first overlap verifications 370 and first extensions 380 are made iteratively, block after block, in a first extension direction (in width in our example), these first extensions being followed by second delimitations 362, overlap verification 372 and extension 382 made iteratively, block (therefore resulting from the first extension) after block in a second extension direction (in height in our example).

[0105] In such embodiments, the method may comprise a delimitation 360, 362 of a temporary block, obtained by expansion of the block considered (here a current block), according to at least one direction of extension, as a function of the inter-word and / or inter-line spaces as explained above.

[0106] The method may also comprise a verification 370, 372 of an overlap in a first extension direction between the temporary block obtained from a current block and another encompassing block of the image and, when such an overlap exists, an extension 380 of the current block to encompass both its own content element(s) (or alternatively those of its temporary block) and the content element(s) of the overlapped encompassing block (or alternatively those of the temporary block of this block).

[0107] In such embodiments, the inter-word space associated with the extended current block may be modified 390. Thus, in certain embodiments, the inter-word space associated with the extended current block may take the value of the inter-word space of the block which was overlapped in the extension direction (here in width). Thus, if the overlapped block has a character size greater (respectively less) than the character size of the current block, the extended current block will be associated with an inter-word space of a size greater (respectively less) than that of the inter-word space previously associated with the current block. Such embodiments make it possible to adapt the method to groupings of blocks comprising characters of different sizes.

[0108] Similarly, in such embodiments, the inter-line space associated with the block extended current block can be modified 392 . Thus, in certain embodiments, the inter-line space associated with the extended current block can take the value of the inter-line space of the block which was overlapped in the extension direction (here in height), so as to take into account a different character size of the overlapped block, compared to the character size of the current block.

[0109] These inter-word or inter-line space modifications may be optional in certain embodiments. In certain embodiments, according to the second example in particular, the steps of delimitation, verification and extension can be implemented iteratively, as long as an extension with a neighboring box can be carried out.

[0110] For example, in some embodiments (suitable for online word grouping), a box encountered on one side of the image (on the left for example) may be extended in width in the same direction (towards the right for example) as much as possible (by successive extension(s) in the case of the second example) (i.e. until no overlap is detected). Of course, a box located on the right of the image may be similarly extended towards the left.

[0111] In some embodiments, once all the boxes have been extended as much as possible in the same extension direction, the method may include at least one delimitation 362, at least one overlap check 372, and at least one extension 382 in a second extension direction (e.g., in height). Note that because of the extensions, a first box that could not be extended in width before its extension in height may therefore also be extended in width (because the box with which it merges was more extended in width).

[0112] In certain embodiments, the delimitation steps 360, 362 and extension steps 380, 382 can be implemented (iteratively or not) by browsing the analyzed image from a corner of the analyzed image (for example the upper left corner of the image), in a direction of travel similar to that of reading (for example from left to right (or vice versa) and from top to bottom). The first box encountered in this travel can be extended in width in the same direction (towards the right for example) by extension(s) (successive or not) as much as possible (i.e. until no overlap is detected), then the first box can then be extended in height in the same direction (downwards for example) by extension(s) (successive or not) as much as possible (i.e. until no overlap is detected).The steps of delimitation 360, 362, overlap verification 370, 372, extension 380, 382 and (optionally) modification 390, 392 of the sizes of the inter-word and / or inter-line spaces can in certain embodiments be repeated starting from one of the blocks adjacent to the first box (which has possibly been extended).

[0113] The embodiments described above can help to identify in the analyzed image blocks corresponding to different text corpora, potentially corresponding to distinct sentences or paragraphs. In addition, such embodiments can be adapted to distinguish between them paragraphs sharing certain row(s) or column(s) of pixels of the analyzed image (in a "newspaper" type layout).

[0114] In some embodiments, after a first extension in a first extension direction, a second extension in a second extension direction may only be performed conditionally, on certain enclosing blocks, for example, it may only be performed on enclosing blocks resulting from a merger of all enclosing blocks aligned in that first extension direction.

[0115] Such embodiments can help, for example, to merge "in height" only blocks corresponding to entire lines and thus help to gather content elements likely to correspond to a paragraph in a "book" type layout.

[0116] The inter-word space has been considered above as a maximum value, making it possible to identify characters belonging to one or more words in the same paragraph. Alternatively, the inter-word space can be used as a boundary between words (whether or not they belong to the same paragraph), the delimitation and / or the extension in width of block(s) being able to make it possible to delimit distinct words.

[0117] As discussed above, some embodiments of the above method may help to identify in the analyzed image blocks corresponding to different text corpora, such as distinct sentences or paragraphs. Isolating sentences or paragraphs (thus words and phrases having a semantic link in general) may facilitate subsequent character recognition on at least some corpora. This may for example be character recognition implemented by the method 300 or executed after analysis of at least one portion of image by the method 300.

[0118] It is noted that the use of the method of the present application does not limit the character recognition method implemented to a particular technique. Indeed, the method can be used before various text detection methods. One example (among others) of character recognition is optical character recognition (or OCR (for Optical Character Recognition in English)).

[0119] In at least some of its embodiments, the method of the present application can be adapted to an analysis of image(s) having content elements of different character sizes (and even of heterogeneous character size as highlighted above).

[0120] In at least some of its embodiments, the method of the present application can be adapted to an analysis of image(s) comprising elements of content of different fonts respecting typography standards. The method of the present application can also be adapted to an analysis of image(s) comprising content elements of heterogeneous font size, in embodiments making it possible to define (or select), via a user interface for example, a reference font of at least one block of an analyzed image.

[0121] As a result, in at least some of its embodiments, the method of the present application can be adapted to an analysis of numerous official documents (this type of document generally respecting typographic standards) as well as to documents generated via computer applications on the market (such as the office pack for example).

[0122] At least certain embodiments of the method of the present application can help to provide solutions that are simpler to implement than certain solutions of the prior art, based for example on neural networks, and without requiring prior learning.

[0123] The method of the present application also offers the advantage of being economical in terms of consumption of computing resources (storage, CPU), in at least some of its embodiments and therefore of being potentially suitable (in such embodiments) for execution in hardware or software environments offering limited computing resources.

[0124] The simplicity of implementation of the method, as well as its economical nature in terms of computing resources, at least in certain embodiments, can help to make the method compatible with real-time use.

[0125] The method of the present application has been presented above in certain examples in connection with width extensions either to the right or to the left, to correspond to a reading direction of a text. In certain embodiments, a default extension direction can be chosen (for example all the images to be analyzed). In other embodiments, an extension direction can be obtained by reading a configuration file or via a user interface, prior to the analysis of an image or a delimitation of a temporary block.

[0126] The case of a zero angle with respect to a horizontal direction has been considered above. The case of rectangular boxes with a non-zero angle can be easily transposed from the detailed case, by replacing the "horizontal" and "vertical" directions of extension by directions of extension parallel to the sides of the boxes.

Claims

Claims

1. Method for analyzing at least one image, implemented by at least one electronic device, said method comprising: - a delimitation, in an image, of blocks of pixels encompassing graphic objects of the image capable of corresponding to alphanumeric characters;- at least one extension of at least a first of the blocks in a first direction of extension to encompass, in said first block, both a first graphic object already encompassed in said first block and at least a second graphic object of at least a second of said blocks of pixels, said at least one extension in said first direction of extension being implemented when said first block of pixels is spaced in said first direction of extension by less than a first number of pixels from one of said at least one second block of pixels, said first number of pixels taking into account a height of at least one first graphic object of said at least one first block of pixels; - at least temporary storage of said first extended block for the purpose of implementing character recognition on said first extended block.;

2. An analysis method according to claim 1, wherein said extension of said first block in said first extension direction includes several second blocks located consecutively in said first extension direction, two consecutive second blocks being spaced apart from each other by less than a second number of pixels, said second number of pixels taking into account a height of the second consecutive block closest to said first block in said first extension direction

3. A method of analyzing at least one image according to claim 1 or 2 wherein the method comprises at least one extension of said first block in a second extension direction to encompass in said first block both said first and second graphic objects already encompassed in said first block and at least one third graphic object of at least one third block of pixels, said at least one extension in said second extension direction being implemented when said first block of pixels is spaced in a second direction of extension by less than a third number of pixels from one of said at least one third block of pixels, said third number of pixels taking into account a height of at least one graphic object of said at least one first block of pixels

4. An analysis method according to claim 3, wherein said extension of said first block in said second extension direction includes several third blocks located consecutively in said second extension direction, two consecutive third blocks being spaced from each other by less than a third number of pixels, said third number of pixels taking into account a height of the third consecutive block closest to said first block in said first extension direction

5. Method for analyzing at least one image according to claim 3 or 4 where said extension according to said second direction of extension is implemented conditionally, when said first block comprises all the graphic objects detected in said image aligned with said first graphic object according to said first direction of extension.

6. Method for analyzing at least one image according to one of claims 1 to 5 where said method comprises an association with said first graphic object of a first character size.

7. A method of analyzing at least one image according to claim 6 where said first number of pixels corresponds to an inter-word spacing of a first character font for said first character size.

8. A method of analyzing at least one image according to claim 2 wherein said method comprises associating said second graphic object with said second consecutive block closest to a second character size and wherein said second number of pixels corresponds to an inter-word spacing of a second character font for said second character size.

9. A method of analyzing at least one image according to claim 7, combined with claim 3, wherein said third number of pixels corresponds to an inter-line spacing of said first character font for said first character size.

10. A method of analyzing at least one image according to any one of said claims 1 to 9 wherein the method comprises obtaining a designation of said first direction of extension.

Citation Information

Patent Citations

  • System and method for analyzing video content using detected text in video frames

    EP1066577B1

  • Page layout determination of an image undergoing optical character recognition

    US20110222771A1

  • Text segmentation of a document

    US20120102388A1

  • Method and system for preprocessing an image for optical character recognition

    US20120219220A1

  • Image processing system with layout analysis and method of operation thereof

    US20160210507A1