Method and device for content analysis
The method optimizes digital content analysis by selectively applying character recognition to relevant content segments, addressing resource-intensive challenges in existing technologies and enhancing efficiency and scalability.
Patent Information
- Application Number
- FR2023014994
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-27
AI Technical Summary
Existing digital content analysis methods, particularly for multimedia content like videos, are resource-intensive and costly due to the need to process large quantities of images, even with advancements in speed and computing power.
A method that identifies and selects only the most relevant content segments containing textual forms, applying character recognition processing only to these segments, thereby optimizing resource usage and reducing processing time and energy consumption.
This approach minimizes memory requirements, processing time, and energy resources while maintaining effective extraction and transcription of textual information from digital content, making it more efficient and scalable for large-scale content analysis.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for analyzing content Field of invention
[0001] This invention relates to the general field of communications. More specifically, it relates to a method for analyzing digital or multimedia content such as video or text content, images, etc. Prior art
[0002] The explosion of multimedia content makes digital content analysis mechanisms all the more relevant today, particularly with a view to classifying them, organizing them, extracting the information they convey, etc. Optical character recognition (OCR) devices are well-known state-of-the-art content analysis devices. Their objective is to recognize characters incorporated in any image, in order in particular to deduce the texts contained in this image. In the remainder of the text, the terms OCR application or module or optical character recognition application or module will be used interchangeably.
[0003] The uses of OCR technology are numerous. For example, this technology is found in devices used to automatically read addresses on paper mail before it is routed by postal services or in software designed to generate editable text from handwritten or photocopied documents.
[0004] OCR-type devices are also increasingly used to extract textual information from non-textual documents, such as images or videos composed of a succession of images. These contents may in fact include textual forms of various kinds such as subtitles or credits or represent objects on which textual forms appear (for example, on license plates, signs, etc.). The extraction of these textual forms can in particular be used to generate separate subtitle files, index content or as part of surveillance devices.
[0005] Notwithstanding, applying OCR processing to video content involves having to process a significant quantity of images (for example, a two-hour film comprises approximately 180,000 images).
[0006] Thus, even if OCR-type content analysis techniques are making progress in terms of speed of execution and computing power, their implementation can prove costly in terms of computing resources and energy resources, this in a context where more and more content must be analyzed, sometimes in real time.
[0007] Furthermore, an optimization problem may also arise when content processing must be industrialized or carried out on a large scale and where, with equal computing resources, it is required to analyze increasingly long content.
[0008] Object and summary of the invention
[0009] The invention responds in particular to these problems by proposing a method for analyzing at least one content comprising a plurality of consecutive content segments, said method comprising at least one iteration implementing: - a step of identifying content segments integrating at least one textual form among said plurality of content segments, - a step of selecting, for at least one textual form, a content segment from among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - a step of triggering an application of character recognition processing on said at least one selected content segment.
[0010] For the purposes of the invention, the term "content segments" means sub-parts or subdivisions of a content. Thus, by way of illustration, the content segments of a video content designate, for example, the images of which the video content is composed; if the content is a slideshow, the content segments designate, for example, the slides of the slideshow; if the content is a digital text file (for example, a .pdf or .doc file), the content segments may designate the pages of this file, etc.
[0011] The invention therefore advantageously proposes to select / sort the most relevant segments of the analyzed content and to limit the application of expensive character recognition processing (for example, OCR processing or character recognition processing based on artificial intelligence, etc.) to these content segments. This makes it possible to minimize the memory required, the processing time and the energy resources required for content analysis, and in particular character recognition.
[0012] The method therefore applies two successive filters to the analyzed content segments.
[0013] More particularly, in a first step, the method focuses on the content segments which integrate at least one textual form (and thus discards / filters the segments which do not integrate a textual form).
[0014] The term “text form” here designates graphic elements corresponding to text (i.e. comprising one or more characters such as letters, numbers, special characters, etc.) but whose characters are not yet necessarily recognized, translated or identified.
[0015] In a second step, the method filters out from among the remaining segments, the irrelevant content segments such as the segments containing redundant textual forms (i.e. presenting similarities with regard to a determined criterion).
[0016] This method therefore advantageously makes it possible to optimize the implementation of the character recognition processing, without loss of information at the level of the texts extracted from the analyzed content.
[0017] According to a particular embodiment, during the selection step, a content segment is selected if it integrates a textual form further verifying a given visibility criterion.
[0018] This embodiment advantageously makes it possible to avoid triggering costly character recognition processing on a content segment incorporating a form that will not be visible or will be barely visible with regard to a given criterion. This criterion may, for example, be attached to the persistence of the textual form over a minimum number of successive content segments. In this way, a content segment incorporating a form that only appears briefly over a reduced number of consecutive content segments (for example, in the manner of a subliminal textual form) will not be selected. According to other examples, such a visibility criterion may refer to a minimum contrast level or a minimum size of the textual form within the content segment.
[0019] According to one embodiment, the similarity criterion relates to a deformation of the textual form over several consecutive content segments.
[0020] This embodiment advantageously makes it possible to avoid triggering several character recognition processes for segments of content having globally similar textual forms although distorted, for example due to a perspective effect or a magnification or reduction effect. This embodiment is of great interest when the analyzed contents include elements on which text is applied, for example license plates on a surveillance video or textures of three-dimensional objects comprising a textual form in a video game.
[0021] The similarity criterion may in particular relate to a partial disappearance of the textual form over several consecutive content segments.
[0022] It is thus possible to avoid applying several character recognition processes on segments of content containing globally similar forms although having parts truncated in relation to each other. This case of This figure can be encountered, for example, in the case of scrolling text or text gradually moving out of the frame of an image. This embodiment is therefore particularly advantageous for reducing the resources required to analyze content containing scrolling text banners.
[0023] According to a particular embodiment of the invention, said at least one content comprises video content and / or a video game and / or text content and / or multimedia content.
[0024] Here, the term “video game” means the video game as it appears and is rendered on a screen and not its computer code.
[0025] The implementation of the analysis method for video-type content or for video games is particularly relevant due, on the one hand, to the fact that this content comprises a very large number of content segments, a significant portion of which do not include a textual form, and on the other hand, that the consecutive content segments including a textual form (for example subtitles, filmed texts, information banners, etc.) have significant similarities. This embodiment therefore makes it possible to substantially optimize the resources mobilized by the character recognition processing application.
[0026] According to one embodiment, said at least one textual form, during the identification step, and / or said similar textual form, during the selection step, is located in a given area of the content segments.
[0027] The identification and consideration of this area makes it possible to reduce the complexity of the processing carried out during the analysis of the content by limiting the size of the segments to be analyzed during the identification and selection steps: it is sufficient, for example, to process only portions of images or slides rather than complete content segments (i.e., complete images or slides). This embodiment therefore makes it possible to limit the resources necessary for the preliminary analysis of the content segments (i.e., the two filtering steps implemented by the invention) while maintaining optimal extraction and transcription of the textual forms integrated into the content.
[0028] According to one embodiment, during the triggering step, the application of the character recognition processing is triggered on at least one segment of content selected in at least said zone.
[0029] By subjecting reduced areas of content segments (e.g., a portion of an image or slide) to character recognition processing rather than complete content segments (e.g., an entire image or slide), it is possible to substantially reduce the complexity of the processing performed.
[0030] In an embodiment in which the analysis method comprises a plurality of iterations, during an iteration, the identification step is only applied to a sub- part of the plurality of content segments, said sub-part being a function of at least one observation data collected during a previous iteration and / or a type of said content.
[0031] This embodiment proposes to exploit data collected during iterations of the analysis method applied to the same content in order to refine the implementation parameters of said method. This can lead to a substantial reduction in the number of segments to be analyzed during the identification and selection steps. This reduction is achieved while maintaining a qualitative analysis of the content.
[0032] Said at least one observation data may comprise in particular: - at least one information representative of a number of consecutive content segments presenting said similar textual form, and / or - at least one piece of information representative of a given area of content segments in which said similar textual form is likely to be found. and / or - at least one piece of information representing an evolution of a textual form integrated into at least one segment of content, for example in the case of an update of part of the content.
[0033] The first type of information typically makes it possible to draw lessons from the second step of the method (i.e. the selection of content segments from among consecutive identified content segments integrating a similar textual form) in order to apply the first step of the method (i.e. the identification of textual form) only to a number of content segments itself restricted or to restricted areas of content segments.
[0034] For example, observation data may be representative of the fact that the subtitles of a video are always displayed on a minimum number of 30 consecutive images. The integration of this observation data causes the method to identify textual forms within an image every 30 images rather than across all the images of the content.
[0035] The second type of observation data mentioned above (which takes into account a given area in which the text forms sought are likely to be found) makes it possible to adapt and optimize the consumption of the resources mobilized by the method according to the types of content analyzed. Certain types of content may have similarities between them. For example, slideshows may contain redundant and / or irrelevant information at the bottom of the page (page number, legal notice). Thus, when the method analyzes content of the presentation / slideshow type, by default and due to the type of content analyzed (here a slideshow), text shape identification can be performed only on the top area of each content segment (i.e. each slide).
[0036] Of course, other observation data can be considered in the context of the invention.
[0037] The invention also relates to a device for analyzing at least one content comprising a plurality of consecutive content segments, said device comprising a plurality of modules activated during at least one iteration, said modules comprising: - an identification module configured to identify content segments integrating at least one textual form among said plurality of content segments, - a selection module configured to select, for at least one textual form, a content segment among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - a trigger module, configured to trigger an application of character recognition processing on the selected identified content segments.
[0038] No limitation is attached to the character recognition processing considered. It may be for example an OCR processing, a character recognition processing based on artificial intelligence, etc.
[0039] The invention also relates to a computer program comprising program code instructions for implementing the content analysis method according to any one of the particular embodiments described above, when this program is executed by a processor.
[0040] Such instructions can be stored permanently in a non-transitory memory medium of a communication terminal implementing the content analysis method according to the invention.
[0041] This program may use any programming language, and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0042] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above.
[0043] The recording medium may be any entity or device capable of storing the program. For example, the medium may include a storage means, such as a ROM (Read Only Memory), for example a CD ROM (Compact Disc Read-Only Memory) or a chip ROM. microelectronics, or a magnetic recording medium, for example a mobile medium, a hard disk or an SSD (“Solid State Drive” in English).
[0044] Furthermore, the recording medium may be a transmissible medium, such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program contained therein is remotely executable. The program according to the invention may in particular be downloaded over a network, for example an Internet-type network.
[0045] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the aforementioned analysis method.
[0046] The invention also relates to a content analysis system, said system comprising: - an analysis device according to the invention, and - a character recognition module, for example OCR type, or based on artificial intelligence previously trained for this purpose.
[0047] The analysis device, the computer program, the recording medium and the analysis system according to the invention have the same advantages mentioned above as the analysis method.
[0048] Furthermore, it is possible to envisage, in other embodiments, that the analysis method, the analysis device and the analysis system according to the invention have in combination all or part of the aforementioned characteristics. Brief description of the drawings
[0049] Other characteristics and advantages will appear on reading particular embodiments of the invention, given as illustrative and non-limiting examples, and the appended drawings, among which:
[0050] [Fig.l] represents a content analysis system according to the invention in a particular embodiment,
[0051] [Fig.2] represents the steps of the analysis method according to a particular mode of realization,
[0052] [Fig.3] illustrates the application of the analysis method according to the invention to a first example of content segments,
[0053] [Fig.4a] and [Fig.4b] illustrate the application of the analysis method according to the invention. to a second example of content segments,
[0054] [Fig.5] illustrates the application of the analysis method according to the invention to a third example of content segments,
[0055] [Fig.6] illustrates the application of the analysis method according to the invention to a fourth example of content segments,
[0056] [Fig.7] illustrates the application of the analysis method according to the invention to a fifth example of content segments,
[0057] [Fig.8] represents an example of a content analysis device according to an embodiment of the invention. Detailed description
[0058] Description of an example of architecture in which the content analysis method is implemented
[0059] Figure [Fig.l] shows an example of architecture within which the content analysis method according to the invention can be implemented in a particular embodiment. This architecture comprises: - a server S on which is stored at least one digital content C comprising a plurality of consecutive content segments SC, - a MAC content analysis device according to the invention, configured to implement the content analysis method according to the invention, - a character recognition module such as, for example, in the embodiment described here, an optical character recognition OCR module, this module being able to be software or hardware. In the example envisaged here, it is a software module and more particularly an optical character recognition OCR application. Alternatively, other character recognition modules can be envisaged, such as for example a module based on a duly trained artificial intelligence, and - a communication network R allowing the server S, the content analysis device MAC and the optical character recognition application OCR to exchange various data (e.g. content and / or content segments) between them, in particular within the framework of the invention.
[0060] In an alternative embodiment, the server C, the MAC content analysis device and the character recognition module are co-located.
[0061] In the embodiment described here, the digital content C designates video content. However, the invention applies to other types of digital content such as a digital presentation (slideshow) composed of slides, a video game, text content, multimedia content, etc.
[0062] According to the embodiments, the server S contains a multitude of contents Ci , Cn (not shown in the figure) of heterogeneous types (e.g. videos, presentations, text files, etc.).
[0063] In the example of video content contemplated herein, the plurality of content segments SC designates the images or frames of the video content. In the In the case of other types of content such as text files or slideshows, the multitude of content segments refers to a page or a slide respectively.
[0064] The MAC content analysis device and the OCR optical character recognition application (or more generally the character recognition module considered) form a SYS system according to the invention. According to the embodiments, the content analysis device and the optical character recognition application are co-located within the same device or are embedded in two separate devices.
[0065] Description of the steps of the content analysis method according to one embodiment
[0066] Figure [Fig.2] represents the main steps of the content analysis method according to the invention as they are implemented by the MAC content analysis device within an architecture similar or identical to that represented in figure [Fig.l], in a particular embodiment of the invention.
[0067] During an identification step E201, after having received from the server S, via the network R, a content C comprising a plurality of content segments SC, the content analysis device MAC stores the content C in a local memory. Then it identifies, among the plurality of content segments of the content C, those integrating at least one textual form.
[0068] In this embodiment, the identification of at least one textual form is based on a solution such as EAST (described in “an Efficient and Accurate Scene Text Detector”, X Zhou & al., 2017,). However, other techniques for identifying a textual form may be envisaged, such as, for example, differentiable binarization (as described in “DB: Real-time Scene Text Detection with Differentiable Binarization”, M. Liao, & al. 2020, AAAI), convolutional neural networks (as described in “Textboxes: Afast text detector with a single deep neural network”, M Liao & al 2017) or any other technique based, for example, on a neural network, previously trained on a plurality of images to recognize textual parts.It is also possible to use a “pattern matching” algorithm or any other technique for detecting textual form in digital content known to those skilled in the art.
[0069] In the embodiment described here, the data generated by the text identification technique applied to the content segments (detected text areas, dimensions, colors, etc.) are recorded in the local memory of the MAC analysis device. The data thus recorded can be used later, or in order to optimize the identification step in subsequent implementations or even in order to identify updates (for example differences) within already analyzed content and, for example, focus only on these new elements.
[0070] It should be noted that forms identified as text forms may turn out not to be interpretable or translatable into characters later (for example, if they are forms from an unknown alphabet or composed of fictitious characters).
[0071] In the embodiment described here, the identification step E201 is applied to all of the segments of the analyzed content. However, in other embodiments, it may be envisaged to apply the identification step E201 only to a sub-part of the plurality of content segments SC of the content C, for example as a function of pre-established configuration parameters specific to the type of analyzed content.
[0072] For example, for a video game type content C transmitted live, the MAC content analysis device can be configured to only apply the identification step to one image every X consecutive images, X designating an integer (e.g. X=30).
[0073] According to another example, the identification step E201 may only relate to a restricted area of the content segments. By way of illustration, such a restricted area may correspond to a rectangle of given dimensions (e.g. 1280 pixels by 200 pixels), placed at the bottom of the image. The information relating to this restricted area (its existence, its dimensions and its coordinates) may be established beforehand for content of the “subtitled films” type.
[0074] Step E201 leads to a set of segments integrating all of the textual forms detected in the content C. This set of segments obtained at the end of step E201 is stored in the local memory of the MAC analysis device. The following step aims to filter the redundant textual forms from this set.
[0075] More precisely during a selection step E202a, the content analysis device MAC selects, for at least one textual form detected during step E201, a content segment from among consecutive content segments identified during the identification step E201 and integrating a textual form similar to said at least one textual form, with regard to a given similarity criterion.
[0076] In this embodiment, the selection step E202a is carried out by means of a solution based on a measurement of the structural similarity index (“structural similarity index measure” in English) applied to the textual forms integrated into the consecutive segments. If the similarity index between the textual forms integrated into two consecutive content segments, measured by the solution, is above a given threshold, the textual forms of the two content segments are considered to be similar, otherwise they are considered to be different or not similar. The invention may, according to another embodiment, implement implement solutions based on other measures such as PSNR (Peak signal-to-noise ratio in English or peak signal-to-noise ratio in French), MSE (Mean squared error in English or Mean quadratic error in French), analyses of content segments by image component (color or mask), DCT (discrete cosine transform) type tools which make it possible to integrate structural elements into the comparison of textual forms, or any other solution, capable of determining whether at least two images are similar based on the calculation of a similarity index of these said images, known to those skilled in the art.
[0077] Step E202a therefore consists of selecting, for each detected textual form, a content segment which is representative of the consecutive content segments integrating a similar textual form.
[0078] For example, if a text form corresponding to a specific subtitle appears on 48 consecutive segments (images) of a video, the method selects one of these segments and discards (i.e., does not select or filter) the other 47. The selected segment is not necessarily the first among the multitude of consecutive segments presenting the text form in question; it may be a segment integrating the text form with the most detail or result from a choice based on another pattern (random selection, segment in the middle of the multitude of consecutive segments, etc.).
[0079] The similarity criterion thus defines the degree of tolerance allowed for considering or judging that two textual forms are similar.
[0080] According to an alternative embodiment, two text forms are considered similar if they are identical.
[0081] According to another embodiment variant, the similarity criterion admits that two consecutive segments integrate a similar textual form even if the form presents a deformation between these two content segments since this deformation corresponds to a deformation admitted by the similarity criterion (for example, homothety, rotation, non-coaxial deformation, truncation / partial disappearance etc.) and therefore considered to be of little significance. Conversely, if the deformation between two textual forms integrated into two consecutive segments is greater than a predefined criterion, the two segments are ultimately considered as integrating two distinct forms.
[0082] For example, two images incorporating a similarly shaped subtitle, present in two distinct scenes of a film, may be selected because the two content segments (images) are not considered duplicates but two segments incorporating a similar textual form in distinct contexts (non-consecutive segments) of the content.
[0083] In another embodiment, during the selection step E202a, the MAC analysis device selects a content segment from among consecutive content segments identified during the identification step E201 integrating a similar textual form, this textual form further verifying a given visibility criterion.
[0084] Different visibility criteria may be envisaged. For example, the visibility criterion may imply that a text form appears on at least a given number of consecutive content segments (persistence), or that the text form is sufficiently large (for example, of a dimension greater than a given number of pixels) or that the text form is sufficiently visible (for example, because it has contrast or brightness characteristics greater than a given criterion) so that a segment incorporating this form is selected.
[0085] The content segments selected during step E202a are stored in the local memory of the MAC content analysis device.
[0086] In E202b, the MAC content analysis device may collect observation data for use in future implementations of the method or not collect such data (O in the figure).
[0087] The collected observation data may relate, for example, to the number of consecutive content segments incorporating the same textual form (for example, the number of images) or to characteristics linked to the determination of the presence of text within portions / zones where the textual forms detected during step E201 are located in the content segments. Such characteristics may, for example, correspond to a zone defined by xy coordinates, etc.).
[0088] These examples are however not limiting in themselves. Thus, one can also envisage observation data corresponding to any type of metadata relating to the textual forms to be generated for example by means of stabilized statistical models, evolutionary models using neural networks or any other data analysis model known to those skilled in the art.
[0089] According to the embodiments, before being used in an identification step, tests and verifications are carried out to evaluate the relevance of the observation data.
[0090] Thus, at the end of step E202a, the content analysis device MAC has a set E of content segments selected from among the plurality of content segments forming the video content C. During a triggering step E203, the content analysis device MAC then triggers the application of an optical character recognition processing by the OCR application on the segments of the set E thus selected in step E202. For this purpose, it sends a command to the OCR application indicating to it the set E of segments on which to carry out the OCR processing.
[0091] In a particular embodiment, the content analysis device MAC receives at E201 only a part / fraction of the segments of the content C to be analyzed (for example the first 100 images of the video content C, or for presentation-type content, the first 20 slides of the presentation, etc.). In this embodiment, it can be envisaged that the MAC analysis device executes the content analysis method iteratively, each iteration corresponding to the analysis of a different part of the segments of the content C to be analyzed and comprising the steps E201, E202 and E203 described previously. Iterations are implemented by the MAC analysis device until it has received and processed all of the segments of the content C.
[0092] In this embodiment where the analysis method is implemented iteratively, during an iteration, the content analysis device MAC can collect during step E202b, observation data such as described previously. These observation data can be advantageously used during step E201 of a subsequent iteration, in a similar or identical manner to what was described previously for two executions of the method, so as to reduce the number of content segments on which step E201 is applied: step E201 is then applied only to a sub-part of the plurality of content segments of the content C supposed to be processed during this iteration.
[0093] Description of the implementation of the invention for particular examples of content segments
[0094] We will now illustrate on different examples of content segments how the analysis method according to the invention is applied,
[0095] First example
[0096] Figure [Fig.3] represents three groups of content segments G3a, G3b and G3c corresponding respectively to groups of content segments obtained respectively upstream of steps E201, E202 and E203 described in figure [Fig.2].
[0097] Thus, group G3a comprises 5 consecutive content segments numbered S1 to S5 received by the MAC analysis device from server S.
[0098] Group G3b corresponds to the content segments integrating at least one textual form, identified among the content segments of group G3 during the identification step E201.
[0099] In this example, only the content segments S2, S3 and S5 have been identified during step E201 and are analyzed during the selection step E202a.
[0100] The group G3c corresponds to the content segments selected during the selection step E202a from among the consecutive identified content segments integrating a similar textual form.
[0101] In this example, the analysis method determines that the textual form integrated into segments S2 and S3 is similar, and selects only one of the two, segment S5 being the only segment integrating this textual form it is also selected.
[0102] Only segments S2 and S5 are targeted by the OCR processing triggering step.
[0103] Second example
[0104] Figures [Fig.4a], [Fig.4b] represent three groups of content segments G4a, G4b and G4c, respectively G4a', G4b' and G4c', corresponding to groups of content segments obtained respectively upstream of steps E201, E202 and E203 described in figure [Fig.2] and iterated a first time [Fig.4a] then a second time [Fig.4b].
[0105] Thus the group G4a comprises 72 consecutive content segments numbered S1 to S72 received by the MAC analysis device from the server S.
[0106] The content segments S2 to S23, S26 to S47, S50 to S71 which are represented by the ellipses incorporate text forms similar to those incorporated respectively in the segments S1 and S24, S25 and S48, S49 and S72.
[0107] Group G4b corresponds to the content segments integrating at least one textual form, identified among the content segments of group G4a, i.e. during the identification step.
[0108] In this embodiment, all the content segments integrate a textual form, therefore all the segments have been identified during step E201 and are analyzed during the selection step E202a.
[0109] The group G4c corresponds to the content segments selected during the selection step E202a from among the consecutive identified content segments integrating a similar textual form.
[0110] In this example, only segments S1, S25 and S49 were selected and, within these segments, only a particular area corresponding to a band at the bottom of the segment was selected and taken into account.
[0111] According to one embodiment, at the end of the OCR processing triggering step, only the areas of these segments are subject to optical character recognition.
[0112] In this embodiment, the observation data relates to the characteristics of an area of interest and the frequency of segments integrating a textual form to be taken into account.
[0113] With reference to Figure [Fig.4b], the group G4a' corresponds to 72 consecutive content segments numbered S73 to S144, which follow the segments SI to S72 of the group G4a of Figure [Fig.4a]. The content segments S74 to S95, S98 to SI 19 and S122 to S143 which are represented by the ellipses incorporate a form textual similar to those integrated respectively by segments S73 and S96, S97 and S120, S121 and S144.
[0114] The group G4b' corresponds to the content segments integrating at least one textual form, identified among the content segments of the group G4a', that is to say during the identification step.
[0115] In this example, the identification step does not process all of the segments of the plurality of content segments but only the segments indicated by the observation data collected during the previous iteration (in the example the bottom bands of the segments and only one segment every 24 segments).
[0116] The identification step therefore only concerns the content segments S73, S97 and S121.
[0117] The G4c' group corresponds to the content segments selected from the consecutive identified content segments integrating a similar textual form.
[0118] In this embodiment, segments S73, S97 and S121 are selected and within these segments, only a particular area corresponding to a band at the bottom of the segment is selected. Only the segment areas are targeted by the OCR processing triggering step.
[0119] Third example
[0120] Figure [Fig.5] describes three groups of content segments G5a, G5b and G5c corresponding to groups of content segments obtained respectively upstream of steps E201, E202 and E203 described in figure [Fig.2] in an embodiment where the visibility criterion, used during the selection step, implies that the textual form or a similar textual form is present on at least 2 consecutive content segments.
[0121] Thus the group G5a comprises 7 consecutive content segments numbered SI to S7 received by the MAC analysis device from the server S.
[0122] Group G5b corresponds to the content segments integrating at least one textual form, identified among the content segments of group G5a, during the identification step E201.
[0123] In this example, all the content segments identified were identified during step E201 and are analyzed during the selection step E202a.
[0124] The group G5c corresponds to the content segments selected during the selection step E202a from among the consecutive identified content segments integrating a similar textual form and for which the textual form satisfies a given visibility criterion. This criterion relates here to the persistence of the textual form over at least a given number of consecutive content segments.
[0125] Consequently, only segments S1 and S5 are selected and are targeted by the OCR processing triggering step.
[0126] Fourth example
[0127] Figure [Fig.6] describes three groups of content segments G6a, G6b and G6c corresponding to groups of content segments obtained respectively upstream of steps E201, E202 and E203 described in figure [Fig.2] in an embodiment where the similarity criterion, used for the selection step, admits that the textual form presents a deformation as long as this is less than a given deformation criterion between 2 consecutive content segments.
[0128] Thus group G6a corresponds to 7 consecutive content segments numbered SI to S7.
[0129] Group G6b corresponds to the content segments integrating at least one textual form, identified among the content segments of group G6a, i.e. at the end of the identification step.
[0130] In this example, only the content segments SI to S6 are identified during step E201 and are analyzed during the selection step E202a.
[0131] Group G6c corresponds to the content segments selected from among the consecutive identified content segments incorporating a similar textual form with respect to a similarity criterion, despite a given degree of deformation of the textual form between consecutive content segments.
[0132] Consequently, during the identification step, two textual forms are identified: a form present on the content segments SI to S5 and a form on the segment S6 (in fact the same form as on the segments SI to S5 but with a deformation, here a rotation, too significant to consider that the two textual forms are similar).
[0133] Only these segments are selected and are targeted by the OCR processing triggering step.
[0134] Fifth example
[0135] Figure [Fig.7] represents three groups of content segments G7a, G7b and G7c corresponding to groups of content segments obtained respectively upstream of steps E201, E202 and E203 described in figure [Fig.2], in an example where the similarity criterion, used for the selection step, admits that the textual form can be affected by a disappearance (a masking) between at least two consecutive content segments according to a given degree (for example if the disappearance is less than a fraction of the textual form or less than a given number of pixels).
[0136] Thus group G7a comprises 7 consecutive content segments numbered SI to S7.
[0137] Group G7b corresponds to the content segments integrating at least one textual form, identified among the content segments of group G7a, during the identification step E201.
[0138] In this example, only the content segments SI to S6 are identified during step E201 and are analyzed during the selection step E202a.
[0139] The group G7c corresponds to the content segments selected during the selection step E202a from among the consecutive identified content segments integrating a similar textual form with regard to the given similarity criterion admitting a partial disappearance (or masking) of the textual form, between at least two consecutive content segments.
[0140] In this example, during the identification step E201, two textual forms are identified: a form present on the content segments S1 to S3 and a form present on the content segments S4 to S6.
[0141] Only segments S1 and S4 are selected. According to the embodiments, the choice of these segments is then made according to the similarity criterion in order to avoid part of the shape present in segments S1 to S3 being also present in segments S4 to S6.
[0142] Only these segments are selected and are targeted by the OCR processing triggering step.
[0143] Description of a content analysis device according to an embodiment of the invention
[0144] [Fig.8] shows the simplified structure of a MAC device configured to implement the content analysis method in a particular embodiment.
[0145] In this embodiment, the MAC device comprises an El / Rl communication module adapted to receive and transmit contents from a server.
[0146] Furthermore, in the particular embodiment of the invention described here, the steps executed by the MAC device, within the framework of the implementation of the content analysis method of the present invention, are implemented by means of instructions of a computer program PG1. For this, the MAC device has the conventional architecture of a computer and notably comprises a memory MEM1, a processing unit UTR1, equipped for example with a processor PROC1, and controlled by the computer program PG1 stored in memory MEM1. The memory MEM1 is a recording medium within the meaning of the invention. The computer program PG1 comprises instructions for implementing the steps of the content analysis method, in particular: - an identification of the content segments integrating at least one textual form among said plurality of content segments, - selecting, for at least one textual form, a content segment from among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - triggering an application of character recognition processing (for example OCR processing) on the selected identified content segments.
[0147] The PG1 program thus defines functional modules of the MAC content analysis device which can be activated for at least one iteration and which include: - an identification module configured to identify content segments integrating at least one textual form among said plurality of content segments, - a selection module configured to select, for at least one textual form, a content segment among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - a trigger module, configured to trigger an application of character recognition processing on the selected identified content segments.
Claims
Claims
1. Method for analyzing at least one content comprising a plurality of consecutive content segments, said method comprising at least one iteration implementing: - a step (E201) of identifying content segments integrating at least one textual form among said plurality of content segments, - a step (E202a) of selecting, for at least one textual form, a content segment among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - a step (E203) of triggering an application of character recognition processing on said at least one selected content segment.
2. Analysis method according to claim 1 comprising a plurality of iterations in which during an iteration, the identification step (E201) is only applied to a sub-part of the plurality of content segments, said sub-part being a function of at least one observation data collected during a previous iteration and / or of a type of said content.
3. Analysis method according to claim 1 or 2 wherein said at least one textual form, during the identification step (E201), and / or said similar textual form, during the selection step (E202a), is located in a given area of the content segments.
4. Analysis method according to claim 3 wherein during the triggering step (E203), the application of the character recognition processing is triggered on at least one selected content segment and in at least said zone.
5. Analysis method according to any one of claims 1 to 4 wherein during the selection step (E202a), a content segment is selected if it integrates a textual form verifying a given visibility criterion.
6. Analysis method according to any one of claims 1 to 5 in which the similarity criterion relates to a deformation of the textual form over several consecutive content segments.
7. Method of analysis according to any one of claims 1 to 6 in which the similarity criterion relates to a disappearance partial text form over several consecutive content segments.
8. Analysis method according to claim 2 wherein said at least one observation data item comprises: - at least one information item representative of a number of consecutive content segments having a similar textual form with regard to the similarity criterion, and / or - at least one information item representative of a given area of content segments in which a similar textual form with regard to the similarity criterion is likely to be found and / or - at least one information item representative of an evolution of a textual form integrated into at least one content segment.
9. Analysis method according to any one of claims 1 to 8 wherein said at least one content comprises video content and / or a video game and / or text content and / or multimedia content.
10. Device (MAC) for analyzing at least one content comprising a plurality of consecutive content segments, said device comprising a plurality of modules activated during at least one iteration, said modules comprising: - an identification module configured to identify content segments integrating at least one textual form among said plurality of content segments, - a selection module configured to select, for at least one textual form, a content segment among consecutive identified content segments integrating a textual form similar to said at least one textual form with regard to a given similarity criterion, and - a trigger module, configured to trigger an application of character recognition processing on the selected identified content segments.
11. Computer program (PG1) comprising program code instructions for implementing an analysis method according to any one of claims 1 to 9, when said program is executed on a computer.
12. Computer-readable information or recording medium (MEM) on which a computer program according to claim 11 is recorded.
13. Content analysis system (SYS), said system comprising: - an analysis device (MAC) according to claim 10, and - a character recognition module (OCR).
Citation Information
Patent Citations
Method and apparatus for recognizing subtitle region, device, and storage medium
US20230027412A1