Image recognition method and device

By recognizing the coordinates of each line of characters in an image and dividing it into paragraphs based on the average line height, the problem of inflexible voice playback in traditional OCR technology is solved, enabling visually impaired users to accurately understand text.

CN120894784APending Publication Date: 2025-11-04VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511011886.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional OCR technology lacks flexibility in recognizing text in images, resulting in fixed voice broadcasting methods that cannot accurately convey text content, especially for visually impaired users.

Method used

By recognizing the coordinate information of each line of characters in an image, combining it with the average line height of the characters to divide the text into paragraphs, and playing the text based on the paragraphs, flexible voice broadcasting is achieved.

Benefits of technology

It improves the flexibility of electronic devices in recognizing image characters, enabling visually impaired users to accurately understand the text content of images and enhances the effectiveness of voice broadcasting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894784A_ABST
    Figure CN120894784A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition method and device, and belongs to the technical field of electronics. The method comprises the following steps: identifying a first image to obtain coordinate information of each row of characters in the first image; dividing the characters in the first image into at least one paragraph based on the coordinate information and the average row height of the characters in the first image; and playing characters in the first image based on the at least one paragraph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, specifically relating to an image recognition method and apparatus. Background Technology

[0002] Currently, Optical Character Recognition (OCR) technology has become an important tool for recognizing text in images. Electronic devices can use OCR technology to convert text in images into computer-readable text. For some users, such as visually impaired users, it is necessary for electronic devices to read aloud the text recognized in images. In this case, the user can trigger the electronic device to read aloud the recognized text in the image in the form of speech, so that visually impaired users can understand the text content in the image through hearing.

[0003] However, because traditional OCR technology recognizes text in images in a fixed way, the recognized text can only be displayed in a fixed way. As a result, the voice broadcast can only be broadcast in a fixed way, which may not allow users to accurately understand the content of the text through hearing. Therefore, the traditional method of recognizing text is not flexible enough and the effect of broadcasting text to users is poor. Summary of the Invention

[0004] The purpose of this application is to provide an image recognition method and apparatus that can improve the flexibility of electronic devices in recognizing characters in images, thereby making the text read aloud to users more effectively.

[0005] In a first aspect, embodiments of this application provide an image recognition method, which includes: recognizing a first image to obtain coordinate information of each line of characters in the first image; dividing the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image; and playing the text in the first image based on the at least one paragraph.

[0006] Secondly, embodiments of this application provide an image recognition device, which includes: a processing module and a playback module; the processing module is used to recognize a first image and obtain coordinate information of each line of characters in the first image; the processing module is also used to divide the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image; the playback module is used to play the text in the first image based on the at least one paragraph obtained by the processing module.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the image recognition method as described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the image recognition method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the image recognition method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the image recognition method as described in the first aspect.

[0011] In this embodiment, the electronic device recognizes a first image to obtain the coordinate information of each line of characters in the first image. Then, based on the coordinate information and the average line height of the characters in the first image, the characters in the first image are divided into at least one paragraph, thereby playing the text in the first image based on at least one paragraph. In this way, when recognizing an image, the electronic device can determine the position of each line of characters based on the coordinate information of each line, and then determine whether there are paragraphs between lines based on the line height of the characters at each position. This allows the characters to be divided into paragraphs, enabling characters within the same paragraph to be grouped together. Furthermore, during voice playback, a paragraph of text can be played continuously, allowing the user to accurately understand the text content in the image through hearing. This improves the flexibility of the electronic device in recognizing characters in images and provides a better reading experience for the user. Attached Figure Description

[0012] Figure 1 This is one of the flowcharts of the image recognition method provided in the embodiments of this application;

[0013] Figure 2 This is a schematic diagram of an example of characters in an image provided in an embodiment of this application;

[0014] Figure 3 This is the second flowchart of the image recognition method provided in the embodiments of this application;

[0015] Figure 4 This is one of the schematic diagrams of the color block marker characters provided in the embodiments of this application;

[0016] Figure 5 This is a second schematic diagram of the color block marker characters provided in the embodiments of this application;

[0017] Figure 6 This is the third schematic diagram of the color block marker characters provided in the embodiments of this application;

[0018] Figure 7 This is a schematic diagram of the four corners of the color block provided in the embodiments of this application;

[0019] Figure 8 This is the third flowchart of the image recognition method provided in the embodiments of this application;

[0020] Figure 9 This is the fourth flowchart of the image recognition method provided in the embodiments of this application;

[0021] Figure 10 This is the fifth flowchart of the image recognition method provided in the embodiments of this application;

[0022] Figure 11 This is the sixth flowchart of the image recognition method provided in the embodiments of this application;

[0023] Figure 12 This is a schematic diagram of the image recognition device provided in the embodiments of this application;

[0024] Figure 13 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0025] Figure 14 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0028] The terms "at least one," "at least one," etc., in this application refer to any one, any two, or a combination of two or more of the included objects. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more, and its meaning is similar to that of "at least one."

[0029] The image recognition method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0030] This application's embodiments can be applied to scenarios involving image recognition. Specifically, in this application's embodiments, the electronic device can recognize characters in an image, including text and symbols.

[0031] For example, in this application embodiment, for certain users, such as visually impaired users or users with reading and writing difficulties (hereinafter referred to as reading-disabled users), after the electronic device recognizes the characters in the image, it can broadcast the recognized text content by voice so that these users can understand the text content in the image through hearing.

[0032] The image recognition method provided in this application will be illustrated below using some specific scenarios as examples.

[0033] Scenario 1: When a dyslexic user wants to access the text content of a paper document, the user can use their mobile phone to take a picture of the document. The phone can then recognize the text content in the photo, aggregate it into paragraphs, and read out the text content of each paragraph in sequence.

[0034] Scenario 2: When a user wants to read aloud the text content of a text image in their phone's photo album, the user can use their phone to recognize the text content in the text image, aggregate the recognized text content into paragraphs, and then play the text content of the selected paragraph based on the user's choice.

[0035] It should be noted that the above scenarios 1 and 2 are merely exemplary examples of some scenarios that may be applied to the embodiments of this application. In actual implementation, the embodiments of this application can also be applied to any possible scenarios that assist users in obtaining text content in text images. The embodiments of this application are not limited here.

[0036] This application provides an image recognition method and apparatus. An electronic device recognizes a first image, obtaining the coordinate information of each line of characters in the first image. Then, based on the coordinate information and the average line height of the characters in the first image, the characters in the first image are divided into at least one paragraph. Based on this, the text in the first image is played based on at least one paragraph. Thus, when recognizing an image, the electronic device can determine the position of each line of characters based on the coordinate information of each line, and then determine whether there are paragraphs between lines based on the line height of the characters at each position. This allows the characters to be divided into paragraphs, enabling characters within the same paragraph to be grouped together. Furthermore, during voice playback, a paragraph of text can be played continuously, allowing the user to accurately understand the text content in the image through hearing. This improves the flexibility of the electronic device in recognizing characters in images and provides a better effect for playing text to the user.

[0037] The image recognition method provided in this application can be executed by an image recognition device, which can be an electronic device, or a functional module or functional entity within an electronic device. The following description uses an electronic device as an example to illustrate the technical solution provided in this application.

[0038] Figure 1 A flowchart of an image recognition method provided in an embodiment of this application is shown, such as... Figure 1 As shown, the image recognition method provided in this application embodiment may include the following steps 201 to 203.

[0039] Step 201: The electronic device recognizes the first image and obtains the coordinate information of each line of characters in the first image.

[0040] Optionally, in this embodiment, the first image includes characters, which may include text, or characters and symbols. The symbols may include, but are not limited to, at least one of punctuation marks and special characters; for example, special characters may be operators. Of course, the characters in the first image may also be in other possible forms, and this embodiment does not impose specific limitations.

[0041] It should be noted that the first image may also include other image elements, such as figures, objects, etc. The specific elements are determined based on actual usage requirements, and this application does not impose any limitations.

[0042] Optionally, in the embodiments of this application, the characters in the first image can be handwritten or printed. The printed characters can be fonts such as KaiTi, SongTi, or LiShu. Of course, the characters in the first image can also be in other possible forms, and the embodiments of this application do not make specific limitations.

[0043] Optionally, in this embodiment, the first image may include, but is not limited to, any of the following: an image captured by the user through the camera of an electronic device, an image downloaded from the Internet, an image obtained by screenshotting, a preview image obtained from any application, or a preview image captured in real time by the camera of an electronic device. Of course, the first image may also be other possible images, and this embodiment does not specifically limit them.

[0044] For example, any of the above applications can be a photo album application, a social application, or a shopping application.

[0045] Optionally, in the embodiments of this application, the first image can be an image in picture format, an image frame in a video, or a moving image.

[0046] In this embodiment of the application, the electronic device can use a first program to recognize a first image and obtain the coordinate information of each line of characters in the first image. The first program has the function of recognizing characters in the image. The first program can be an application program, component, or function program, etc.

[0047] In this embodiment of the application, the electronic device can use OCR technology to recognize the first image to obtain the coordinate information of each line of characters in the first image. The process may include image preprocessing and text region detection.

[0048] During image preprocessing, the electronic device can perform operations such as grayscale conversion, binarization, and tilt correction on the first image, transforming it into a black and white image with horizontally aligned text lines, thereby highlighting the text display area. During text region detection, the electronic device can use projection to locate the region containing each character in the first image and obtain the positional information of each character.

[0049] In this embodiment of the application, the electronic device can first perform image preprocessing on the first image, including grayscale conversion, binarization, tilt correction and other operations, and then use projection to perform text region detection on the obtained black and white image to obtain the coordinate information of each character in the first image.

[0050] In the image preprocessing process, the electronic device performs the following steps: Grayscale conversion calculates the grayscale value of each pixel based on its red, green, and blue color components, converting the first image into a grayscale image. Binarization sets the grayscale value of each pixel to 255 or 0 based on whether it exceeds a threshold, further converting the first image into a black-and-white image where background pixels are white and text pixels are black. The tilt correction detects aligned boundary pixel sets to identify the straight line corresponding to the text line in the first image. Based on the angle between this line and the horizontal direction, the tilt angle of the text is determined, and rotation is used to ensure the text line is horizontal, completing the preprocessing of the first image.

[0051] After preprocessing the first image, the electronic device can use projection to detect text regions. The principle is to locate character coordinates by analyzing the pixel density distribution of characters in the horizontal and vertical directions. First, the electronic device generates a horizontal projection histogram by counting the number of black pixels row by row. Since the number of black pixels in the character row region is significantly higher than the background, it appears as a peak region in the histogram. The Y-axis range of this histogram peak represents the upper and lower boundaries of the character row, while the troughs correspond to the row spacing. Next, vertical projection is performed sequentially within each single-line text region, i.e., the pixel density is counted column by column to obtain the vertical projection histogram corresponding to each character row. Since the number of black pixels in the column containing a single character is significantly higher than the background, the X-axis range of each peak represents the left and right boundaries of the single character corresponding to that peak, while the troughs correspond to the horizontal spacing between characters within each character row. Finally, the Y-axis range corresponding to each line of characters and the X-axis range corresponding to a single character within each line are combined to obtain the coordinate information of each character.

[0052] In this embodiment of the application, the coordinate information mentioned above may include pixel-level coordinates.

[0053] In this embodiment, the coordinate system corresponding to the above coordinate information can be a screen coordinate system, that is, a two-dimensional coordinate system with the top left corner of the screen as the origin, the positive X-axis pointing horizontally to the right, and the positive Y-axis pointing vertically downwards; the coordinate system corresponding to the above coordinate information can also be an image coordinate system, that is, a two-dimensional coordinate system with the top left corner of the first image as the origin, the positive X-axis pointing horizontally to the right, and the positive Y-axis pointing vertically downwards. For example, if the coordinate system corresponding to the above coordinate information is a screen coordinate system, then the pixel-level coordinates (1,2) represent the pixels in the first row from top to bottom and the second column from left to right on the screen of the electronic device. If the coordinate system corresponding to the above coordinate information is an image coordinate system, then the pixel-level coordinates (1,2) represent the pixels in the first row from top to bottom and the second column from left to right on the first image.

[0054] Optionally, in this embodiment of the application, the coordinate information of each row of characters in the first image may include the coordinate information of the first character and the coordinate information of the last character in each row.

[0055] The coordinate information of the first character can include the coordinates of the top left corner and the bottom left corner of the first character, and the coordinate information of the last character can include the coordinates of the top right corner and the bottom right corner of the last character.

[0056] Optionally, in the embodiments of this application, the coordinate information of each row of characters in the first image may also include any of the coordinates of the four corners of each character in each row, the coordinates of the center point of each character in each row, the coordinates of the four corners of the area where each row of characters is located, and other coordinate information. The embodiments of this application do not impose specific limitations.

[0057] For example, such as Figure 2As shown, assume that the characters in the first image 10 include "Time is hasty, yet the ideals of youth will not fade away because of this. We carry the promise of youth and have never stopped moving forward. In the glistening moonlight, I scoop up the clearest cup; in the falling afterglow, I embrace the warmest ray; among the blazing red leaves, I pick the hottest one. The flower is most beautiful when half - open, and the feeling is most profound when left blank. Knowing how to leave blank space for life is also a kind of wisdom in life. In a quiet heart, there is always the most beautiful scenery. Despite the complexity of the world, this heart remains the same, and the feelings remain the same. Some human relationships need to be indifferent to last long. A plain life is more lasting. Only by leaving blank space can there be a long - distance. Keep the lotus in your heart and be adaptable to circumstances. The most beautiful thing in life is the distance of understanding.", the electronic device can recognize the first image, obtain the coordinate information of each character in the first image, and use the coordinate information of the first and last characters in each row of characters in the first image as the coordinate information of the characters in that row. For example, the upper - left corner of the first character "岁" in the first row has coordinates (125, 50) in the screen coordinate system, the lower - left corner has coordinates (125, 70), the upper - right corner of the last character "却" in the first row has coordinates (425, 50) in the screen coordinate system, and the lower - right corner has coordinates (425, 70). Then the coordinate information of the characters in the first row is (125, 50), (125, 70), (425, 50), and (425, 70); the upper - left corner coordinate of the first character "不" in the second row is (75, 80), the lower - left corner coordinate is (75, 100), the upper - right corner coordinate of the last character "青" in the second row is (425, 80), and the lower - right corner coordinate is (425, 100). Then the coordinate information of the characters in the second row is (75, 80), (75, 100), (425, 80), and (425, 100).

[0058] For the coordinate information of the characters in other rows in the first image, the determination method is similar to that of the characters in the first row and the second row above, and will not be listed one by one here.

[0059] Step 202: The electronic device divides the characters in the first image into at least one paragraph based on the coordinate information of each row of characters in the first image and the average line height of the characters in the first image.

[0060] In the embodiment of the present application, dividing the characters in the first image into at least one paragraph means: dividing the characters in the first image into at least one natural paragraph.

[0061] It can be understood that the average line height of the characters in the first image refers to: the average vertical spacing of all character lines in the first image, and the average line height of the characters in the first image can be calculated from the coordinate information. For the calculation of the average line height of characters, reference can be made to the following Step 204a, Step 204b, and Step 204c, or reference can be made to the following Step 204a, Step 204b, and Step 204d. This application will not elaborate here.

[0062] In this embodiment of the application, the row height of all characters in the first image may be the same, or may be different, or some characters in the first image may have the same row height.

[0063] In this embodiment, the electronic device can determine the spacing between rows of characters in the first image based on the coordinate information of each row of characters and the average row height of characters in the first image, thereby dividing the image into paragraphs. For example, the electronic device can use a first color block to mark the area of ​​each row of characters based on the coordinate information of the first and last characters in each row of characters in the first image. Then, based on the average row height of characters in the first image, it can adjust the height of the first color block corresponding to each row of characters to obtain at least one second color block. Finally, it can divide the characters contained in the overlapping second color blocks into the same paragraph to obtain at least one paragraph. For details, please refer to the specific descriptions of steps 202a to 202c in the following embodiments, which will not be repeated here.

[0064] It should be noted that in the embodiments of this application, each line of characters can also be called a character line, which can be understood as the line in which the characters are located.

[0065] Step 203: The electronic device plays the text in the first image based on at least one paragraph.

[0066] In this embodiment of the application, the electronic device can perform text processing and speech synthesis on the text in at least one paragraph to generate a corresponding speech signal and play the text in the first image.

[0067] In this embodiment, the playback of text in the first image can be performed in any possible application. For example, the application can be the application containing the first image, or it can be an application selected by the user.

[0068] For example, the above-mentioned application can be a music player application or a social application.

[0069] This application provides an image recognition method. When recognizing an image, the electronic device can determine the position of each character based on the coordinate information of each line of characters. Then, based on the line height of each character, it can determine whether there is a paragraph between the lines, thereby dividing the characters into paragraphs. This allows characters within the same paragraph to be grouped together, and thus, during voice broadcasting, the text content of a natural paragraph can be broadcast continuously, enabling users to accurately understand the text content in the image through hearing. This improves the flexibility of the electronic device in recognizing characters in the image and provides a better effect for broadcasting text to users.

[0070] Optionally, in this embodiment, the coordinate information of each row of characters in the first image includes the coordinate information of the first and last characters in each row. For example, combined with... Figure 1 ,like Figure 3 As shown, step 202 can be implemented through steps 202a, 202b and 202c as described below.

[0071] Step 202a: The electronic device uses a first color block to mark the area where each line of characters is located, based on the coordinate information of the first and last characters in each line of characters.

[0072] In this embodiment of the application, for each line of characters in the first image, after the electronic device obtains the coordinate information of the first and last characters in the line, it can determine the position of the first and last characters in the line based on the coordinates of the upper left and lower left corners of the first character and the upper right and lower right corners of the last character in the line. Then, the electronic device can use the first color block to mark from the first character to the last character in the line.

[0073] Optionally, in this embodiment, the color of the first color block can be any color different from the background color of the first image. For example, assuming the background color of the first image is white, the color of the first color block can be red, yellow, or gray. Of course, the first color block can also be other possible colors, and this embodiment does not specifically limit it.

[0074] Optionally, in the embodiments of this application, the shape of the first color block can be any shape. For example, the first color block can be a rectangle or an ellipse. Of course, the first color block can also be other possible shapes. The embodiments of this application do not make specific limitations.

[0075] Optionally, in this embodiment of the application, the height of each first color block can be greater than or equal to the line height of its corresponding character line, and the length of the first color block can be greater than or equal to the length of its corresponding character line.

[0076] For example, combined Figure 2 ,like Figure 4 As shown, the electronic device can use black rectangular color blocks to mark the area where each line of characters in the first image 10 is located. The first image has a total of 16 first color blocks, which are marked as a1 to a16 from top to bottom.

[0077] It should be noted that, in the embodiments of this application, if the first or last character of each line is a punctuation mark or other non-text character, the electronic device can, after subsequently detecting that the character is a non-text character using a convolutional neural network, translate the coordinates of the non-text character so that the rectangles corresponding to the coordinates of the four corners of the character have the same height as the rectangles corresponding to the coordinates of the four corners of the adjacent characters.

[0078] Step 202b: The electronic device adjusts the height of the first color block corresponding to each row of characters based on the average row height of the characters in the first image to obtain at least one second color block.

[0079] In this embodiment of the application, adjusting the height of the first color block corresponding to each line of characters can be done by increasing the height of the first color block.

[0080] In this embodiment of the application, the second color block can be determined as a second color block after adjusting the height of the first color block corresponding to each line of characters.

[0081] In this embodiment, the electronic device can adjust the height of the first color block based on morphological dilation. Specifically, the electronic device can calculate the product of the average line height of the characters in the first image and a preset multiple, and generate a structural element with a length value equal to the product. This structural element can be a vertical line segment or other form of element. The structural element is used to slide in the area where the character line is located to adjust the height of the color block corresponding to the character line. The electronic device can slide the structural element on each first color block and merge all pixels in the area covered by the sliding structural element on each first color block to obtain a second color block corresponding to each first color block.

[0082] For example, the aforementioned preset multiple can be a default value of the electronic device system, such as 0.6. Of course, the aforementioned preset multiple can also be other values, and can be set according to actual needs. This application embodiment does not impose any restrictions.

[0083] For example, combined Figure 4 ,like Figure 5 As shown, the electronic device can calculate the product of the average line height of the characters in the first image and 0.6, and generate a structuring element with a length value of the product. The structuring element can be a vertical line segment. The electronic device can slide the structuring element on color blocks a1 to a16 respectively, and merge all pixels in the area covered by the structuring element on each color block a1 to a16 to obtain the second color block corresponding to each color block a1 to a16. That is, there are 16 second color blocks corresponding to color blocks a1 to a16, and the second color blocks corresponding to color blocks a1 to a16 are color blocks b1 to b16 respectively.

[0084] It should be noted that, in order to more clearly illustrate the second color block, Figure 5 The diagram illustrates each second color block using different filling methods. In actual implementation, the second color blocks can be marked using black rectangular color blocks.

[0085] Step 202c: The electronic device divides the characters contained in at least one overlapping second color block into the same paragraph to obtain at least one paragraph.

[0086] In this embodiment of the application, the overlapping second color blocks among the at least one second color block can be defined as follows: if two adjacent second color blocks overlap, then those two adjacent second color blocks are considered overlapping second color blocks. Therefore, the overlapping second color blocks among the at least one second color block can include at least two second color blocks. For example, if color block b1 overlaps with color block b2, and color block b2 overlaps with color block b3, the electronic device can divide the characters corresponding to color blocks b1, b2, and b3 into one paragraph; if color block b3 does not overlap with color block b4, then the characters corresponding to color blocks b3 and b4 are divided into two different paragraphs.

[0087] The determination of the other second color blocks in the first image is similar to that of color blocks b1 to b4 mentioned above, and will not be listed here.

[0088] It should be noted that the aforementioned overlap refers to a partial overlap between one second color block and another. In the embodiments of this application, overlapping second color blocks are those with partially connected regions.

[0089] It is understandable that since the spacing between paragraphs is greater than the line spacing between character lines within a paragraph, after adjusting the height of the first color block, the line spacing can be covered by the color block while the spacing between paragraphs can be preserved. This allows the second color blocks corresponding to adjacent character lines within a paragraph to partially overlap, while gaps still exist between paragraphs. Thus, the character lines corresponding to the overlapping second color blocks can be divided into the same paragraph, while the character lines corresponding to the non-overlapping adjacent second color blocks (i.e., the second color blocks corresponding to adjacent character lines) are divided into different paragraphs, thereby achieving the purpose of paragraph segmentation.

[0090] In this embodiment, the electronic device can use connected component analysis to detect whether there is overlap between the second color blocks. The electronic device can traverse the image line by line. If a pixel with the same pixel value as the pixel value corresponding to the second color block is detected, the pixel is used as a seed point to start detecting the neighboring pixels of the pixel. If at least one pixel in the four directions of its adjacent pixels has a pixel value equal to the pixel value corresponding to the second color block, they are classified into the same connected component. The pixel with the pixel value equal to the pixel value corresponding to the second color block is used as a new seed point to continue detecting the adjacent pixels until all the pixels in the first image have been traversed. During the traversal, each connected component is assigned a unique identifier. After the traversal is completed, the first image can be divided into at least one connected component according to the unique identifier. Each connected component consists of an independent second color block or at least two overlapping second color blocks. The text in the second color block contained in each connected component is a paragraph.

[0091] In this embodiment of the application, if the first image contains N rows of characters, and each row of characters corresponds to a second color block, the electronic device can iterate through whether there is overlap between any two second color blocks.

[0092] If any two adjacent second color blocks do not overlap, then the characters corresponding to each second color block constitute a paragraph.

[0093] or,

[0094] If some of the second color blocks overlap with other second color blocks, then the text contained in the first image contains at least two paragraphs.

[0095] or,

[0096] If all second color blocks overlap with other second color blocks, then the text contained in the first image is the same paragraph.

[0097] For example, if the second color block a overlaps with the second color block b, and the second color block b overlaps with the second color block c, then the electronic device can divide the text contained in the second color block a, the second color block b, and the second color block c into the same paragraph, which can be called paragraph a. If the second color blocks a to c do not overlap with the second color block d, the electronic device can divide the second color block d into a separate paragraph, which can be called paragraph b.

[0098] In this way, electronic devices can divide the text into paragraphs by morphologically expanding the color blocks, so that subsequent electronic devices can play the text according to the paragraph divisions, thus improving the playback effect.

[0099] Optionally, in the embodiments of this application, step 202c can be implemented by steps 302a to 302e as described below.

[0100] Step 302a: The electronic device identifies at least one overlapping second color patch as a third color patch, thereby obtaining at least one third color patch.

[0101] In this embodiment of the application, determining the overlapping second color block as a third color block means treating all overlapping second color blocks as a whole and taking the resulting color block as the third color block.

[0102] For example, combined Figure 5 ,like Figure 6 As shown, the electronic device can identify overlapping color blocks b1, b2, and b3 as a third color block, called color block c1, and can identify color blocks b4, b5, and b6 as a third color block, called color block c2. For other second color blocks in the first image, similar to the method of identifying color blocks b1 to b6, multiple third color blocks, namely color blocks c1 to c5, can be obtained.

[0103] In this embodiment of the application, the color, shape, or size of each third color block can be determined by all the second color blocks corresponding to that third color block.

[0104] In this embodiment of the application, the electronic device can detect whether there is overlap between the second color blocks by means of connected component detection, thereby determining the overlapping second color blocks as a third color block. The relevant technical principles of connected component detection have been given above and will not be repeated here.

[0105] Step 302b: The electronic device obtains the four boundary information of the contour of each third color block based on the contour coordinate information of each third color block.

[0106] In this embodiment of the application, the contour coordinate information of each third color block can be a set of pixel coordinate points corresponding to the contour of each third color block.

[0107] In this embodiment, the electronic device can first perform grayscale and binarization processing on the image to obtain an image with a white background and a black area where the third color block is located. That is, the grayscale value of the pixels in the background is 255, and the grayscale value of the pixels in the area where the third color block is located is 0. The electronic device can scan the pixels in the image sequentially. If a pixel in the image is detected to have a grayscale value of 0, it checks whether the grayscale values ​​of the four adjacent pixels above, below, left, and right of that pixel are all 0. If the grayscale values ​​of its four neighboring pixels are also all 0, it indicates that the point is an interior point of the third color block; otherwise, it considers the point to be a point on the outline of the third color block and records the pixel-level coordinates of the point. After scanning, the electronic device can obtain the coordinates of all pixels on the outline of the third color block.

[0108] Optionally, in this embodiment of the application, the outline coordinate information of each third color block can also be a set of vector coordinates corresponding to the outline of each third color block.

[0109] Step 302c: The electronic device obtains the coordinate information of the four vertices corresponding to the four boundaries based on the four boundary information of the outline of each third color block.

[0110] In this embodiment, the four boundary information of each third color block refers to the position information of the upper boundary, lower boundary, left boundary, and right boundary of each third color block. The boundary information can be represented by a straight line expression, or by the coordinates of at least two points selected on the straight line. Of course, the boundary information can also adopt other possible forms of expression, which will not be listed one by one in this embodiment.

[0111] In this embodiment, the electronic device can obtain the boundary point coordinates in the outline coordinates of the third color block, such as the coordinates of the leftmost pixel, the rightmost pixel, the topmost pixel, and the bottommost pixel, based on the outline coordinate information of the third color block. Then, the expression of the vertical line passing through the leftmost pixel is used as the left boundary information of the third color block's outline; the expression of the vertical line passing through the rightmost pixel is used as the right boundary information of the third color block's outline; the expression of the horizontal line passing through the topmost pixel is used as the upper boundary information of the third color block's outline; and the expression of the horizontal line passing through the bottommost pixel is used as the lower boundary information of the third color block's outline. Based on the obtained four boundary information, the coordinates of the four vertices obtained by the intersection of the four boundaries are obtained.

[0112] Step 302d: The electronic device marks the border corresponding to each third color block based on the coordinate information of the four vertices.

[0113] In this embodiment of the application, the electronic device can sequentially connect the four vertices of each of the above-mentioned third color blocks to obtain the border of each third color block.

[0114] For example, combined Figure 6 ,like Figure 7 As shown, the four boundaries of color block c1 are s1 to s4, and the four vertices corresponding to the four boundaries of color block c1 are x1 to x4. The electronic device can connect x1 to x4 sequentially to obtain the border corresponding to color block c1. The four boundaries of color block c3 are s5 to s8, and the four vertices corresponding to the four boundaries of color block c3 are x5 to x8. The electronic device can connect x5 to x8 sequentially to obtain the border corresponding to color block c3. It should be noted that the marking method for the borders of each third color block in the first image is similar to that for color blocks c1 and c3, and will not be listed here.

[0115] Step 302e: The electronic device divides the characters contained in each border into the same paragraph to obtain at least one paragraph.

[0116] It can be understood that the border of each third color block contains its corresponding third color block, and the area of ​​each third color block corresponds to the area of ​​each paragraph. Therefore, the characters contained in each border can be divided into the same paragraph.

[0117] For example, combined Figure 7 The text within the border corresponding to color block c1, which reads "Time flies, but the ideals of youth will not fade. We carry the promise to youth and have never stopped moving forward," is grouped into the same paragraph. Similarly, the text within the border corresponding to color block c3, which reads "Flowers are most beautiful when half-open, emotions are most intense when left blank; knowing how to leave blank spaces in life is also a kind of wisdom," is also grouped into the same paragraph. For each color block in the first image, paragraphs are divided using a similar method to color blocks c1 and c3. This results in at least one paragraph being the paragraph contained within the borders corresponding to color blocks c1 through c5.

[0118] In this way, electronic devices can use regular rectangular borders to group the characters contained in each paragraph together to mark each paragraph. This allows the text recognized in the image to be displayed in paragraph form, improving the display effect and making it easier for users to view and understand the text content.

[0119] Optionally, the image recognition method provided in this application embodiment further includes the following steps 204a, 204b and 204c, or the image recognition method provided in this application embodiment further includes the following steps 204a, 204b and 204d.

[0120] Step 204a: The electronic device calculates the line height of each line of characters based on the coordinate information of any character in each line.

[0121] In this embodiment of the application, the coordinate information of any of the above characters may include the upper boundary coordinate information and the lower boundary coordinate information of any character.

[0122] The upper boundary coordinate information of the character can include the coordinates of the upper left corner of the character, and the lower boundary information of the character can include the coordinates of the lower left corner of the character.

[0123] In this embodiment of the application, for the line height of each line of characters in the first image, the electronic device can calculate the distance between the upper left corner coordinate and the lower left corner coordinate of the first character in the line of characters, and determine the distance as the line height of the line of characters.

[0124] For example, combined Figure 2, the electronic device can use the first character "岁" in the first row to calculate the line height of the characters in the first row, that is, subtract the Y value 50 corresponding to the upper left corner coordinate of "岁" from the Y value 70 corresponding to the lower left corner coordinate of "岁", and the obtained value 20 is the line height of the characters in the first row, and its unit is pixels.

[0125] Optionally, in the embodiments of the present application, the upper boundary coordinate information and the lower boundary coordinate information of any of the above characters may further include at least one of the following: the expressions of the upper and lower boundary lines of the characters, and the coordinates of the pixels respectively selected on the upper and lower boundary lines of the characters.

[0126] Step 204b, the electronic device calculates the first average line height based on the line height of each row of characters.

[0127] In the embodiments of the present application, the electronic device can calculate the arithmetic mean, weighted mean or harmonic mean of the line heights of all character lines in the first image to obtain the first average line height.

[0128] Optionally, in the embodiments of the present application, the electronic device can calculate the arithmetic sum of the line heights of all character lines in the first image, and then divide the obtained arithmetic sum by the total number of character lines in the first image to obtain the above first average line height.

[0129] Step 204c, if the ratio of the line height of all row characters in the first image to the first average line height is less than or equal to the first threshold, the electronic device determines the first average line height as the character average line height of the first image.

[0130] In the embodiments of the present application, the conventional value of the above first threshold may be 1.5. Of course, the first threshold may also be other values, which can be specifically set according to actual needs, and the embodiments of the present application do not limit this.

[0131] It can be understood that when the ratio of the line height of all row characters in the first image to the first average line height is less than or equal to the first threshold, it means that there is no character line with too high line height in the first image, that is, the first average line height can better represent the overall line height of the text in the first image. At this time, the first average line height can be used as the character average line height of the first image.

[0132] For example, suppose there are three character lines in the first image, namely character line a, character line b and character line c. The line height of character line a is 20 pixels, the line height of character line b is 30 pixels and the line height of character line c is 10 pixels. Then the first average line height of the first image is 20 pixels. The ratio of the line height of character line a to the first average line height is 1, the ratio of the line height of character line b to the first average line height is 1.5 and the ratio of the line height of character line c to the first average line height is 0.5. Since the ratio of the line height of all character lines in the first image to the first average line height is less than or equal to the first threshold, the electronic device determines the first average line height, i.e., 20 pixels, as the average line height of the characters in the first image.

[0133] Step 204d: If the ratio of the height of at least one row of characters to the first average row height is greater than the first threshold, the electronic device calculates the average row height of the characters in the first image based on the coordinate information of the remaining rows of characters in the first image excluding the at least one row of characters.

[0134] It is understood that when the ratio of the height of at least one line of characters in the first image to the first average line height is greater than the first threshold, it indicates that there are character lines with abnormal line heights in the first image, that is, character lines with excessively high line heights, such as the title line. Character lines with abnormal line heights will affect the overall representativeness of the first average line height. Therefore, the above-mentioned at least one character line with abnormal line heights can be discarded, and the average line height of the remaining lines can be recalculated as the average line height of the characters in the first image.

[0135] In this embodiment of the application, when the ratio of the height of at least one line of characters in the first image to the first average line height is greater than the first threshold, the at least one line of characters can be marked. Then, all character lines in the first image are traversed, and the height of the unmarked character lines is recorded. The electronic device can calculate the arithmetic sum of the heights of all unmarked character lines, and then divide it by the total number of unmarked character lines to determine the obtained value as the average line height of the characters in the first image.

[0136] It should be noted that steps 204a, 204b, and 204c are performed before step 202b. Alternatively, steps 204a, 204b, and 204d are performed before step 202b.

[0137] In this way, the electronic device calculates the average line height of the characters in the first image based on the line height of each line of characters and the average line height of all lines of characters. This can prevent characters with abnormal line heights in the first image from affecting the value of the average line height of the characters in the first image, thereby improving the accuracy of subsequent paragraph segmentation using the average line height.

[0138] Optionally, in this embodiment of the application, the aforementioned at least one paragraph refers to N paragraphs, where N is an integer greater than 1. For example, in combination with... Figure 1 ,like Figure 8 As shown, step 202 can be implemented through steps 205a and 205b as described below.

[0139] Step 205a: The electronic device divides the characters in the first image into M segments based on the coordinate information of each line of characters in the first image and the average line height of the characters in the first image.

[0140] In the embodiments of this application, M is a positive integer.

[0141] In this embodiment, the electronic device can use a first color block to mark the area where each row of characters is located, based on the coordinate information of the first and last characters in each row of characters in the first image. Then, based on the average row height of the characters in the first image, the height of the first color block corresponding to each row of characters is adjusted to obtain at least one second color block. The characters contained in the overlapping second color blocks are then divided into the same segment, resulting in M ​​segments. This division can be understood as the first division. It can be understood that the electronic device can obtain M segments by executing steps 202a to 202c.

[0142] It should be noted that for specific explanations of steps 202a to 202c, please refer to the descriptions in the above embodiments, which will not be repeated here.

[0143] Step 205b: The electronic device divides the characters in the M paragraphs into N paragraphs based on the coordinate information and average width of each line of characters in the M paragraphs.

[0144] In this embodiment of the application, N is greater than M.

[0145] In this embodiment of the application, the above-mentioned N paragraphs can be understood as: the paragraphs obtained by the electronic device from the second division of the M paragraphs obtained after the first division.

[0146] In this embodiment, for each of the M paragraphs, the electronic device can calculate the difference between the coordinates of the first character of each line in a paragraph and the coordinates of the first character of the previous line. Then, it calculates the ratio between this difference and the average width of each line of characters in that paragraph, and so on, to obtain the ratio corresponding to each line of characters in each of the M paragraphs. Therefore, for each of the M paragraphs, the electronic device can determine whether there exists a line in each paragraph whose corresponding ratio is greater than a second threshold. If so, that line and its preceding line are divided into two different paragraphs, thus dividing the characters in the M paragraphs into N paragraphs. It can be understood that the electronic device can divide the characters in the M paragraphs into N paragraphs by executing the following steps 305a and 305b. For a detailed explanation of steps 305a and 305b, please refer to the description in the following embodiments, which will not be repeated here.

[0147] It is understood that, when performing steps 205a and 205b above, step 203 above can specifically be: the electronic device plays the text in the first image based on N paragraphs.

[0148] In this way, electronic devices can use the coordinate information and average width of each line of characters in each paragraph after the initial division to perform a secondary division of the paragraphs obtained after the initial division, thereby improving the accuracy of paragraph division.

[0149] Optionally, in the embodiments of this application, combined with Figure 8 ,like Figure 9 As shown, step 205b can be implemented through steps 305a, 305b and 305c as described below.

[0150] Step 305a: The electronic device calculates the difference between the coordinate information of the first character of each line in the first paragraph and the coordinate information of the first character of the previous line.

[0151] In this embodiment of the application, the first paragraph mentioned above is any one of the M paragraphs.

[0152] In this embodiment, the difference between the coordinates of the first character in each line of the first paragraph and the coordinates of the first character in the previous line can be the difference between the top-left corner coordinate of the first character in each line of the first paragraph and the top-left corner coordinate of the first character in the previous line. It should be noted that this difference can be an algebraic difference, and it can be positive or negative.

[0153] For example, combined Figure 2, the upper left coordinate of the first character '岁' in the first line of characters is (125, 50), and the upper left coordinate of the second character '不' in the second line of characters is (75, 80). The X value 75 corresponding to the upper left coordinate of the character '不' can be subtracted from the X value 125 corresponding to the upper left coordinate of the character '岁', and the obtained -50 is the difference between the coordinate information of the first character in the second line and the coordinate information of the first character in the previous line.

[0154] Step 305b: The electronic device calculates the ratio between the above difference and the average width of each line of characters in the first paragraph.

[0155] In the embodiment of the present application, the average width of each line of characters can be the value obtained by dividing the total length of each line of characters by the total number of characters in each line.

[0156] Among them, the total length of each line can be obtained by calculating the distance between the upper left coordinate of the first character and the upper right coordinate of the last character in each line.

[0157] Exemplarily, in combination with Figure 2 , Figure 2 , the total length of the first line of characters in

[0158] can be obtained by subtracting the X value 125 corresponding to the upper left coordinate of the character '岁' from the X value 425 corresponding to the upper right coordinate of the character '却'.

[0159] In the embodiment of the present application, L is a positive integer.

[0160] It should be noted that for each of the above M paragraphs, the division judgment and division of each of the M paragraphs can be realized by executing Step 305a, Step 305b, and Step 305c to obtain N paragraphs.

[0161] For example, assuming M is 3, the electronic device divides the characters in the first image into three segments, namely segment d1, segment d2, and segment d3, through step 205a. Then, it determines whether there is at least one line in each of the three segments that satisfies the secondary division condition, that is, whether there is at least one line in each of the three segments whose horizontal distance to the average width of the first character and the character above it exceeds a second threshold. Assuming that there is one line in segment d1 that satisfies the secondary division condition, segment d1 is further divided into segment d4 and segment d5. If no line in segment d2 satisfies the secondary division condition, segment d2 is not further divided. If two lines in segment d3 satisfy the secondary division condition, segment d3 is further divided into segment d6, segment d7, and segment d8. That is, through steps 305a and 305b, M segments can be further divided into N segments, which include segment d4, segment d5, segment d6, segment d7, and segment d8.

[0162] In this embodiment of the application, the conventional value of the second threshold can be 1.5. Of course, the second threshold can also be other values, which can be set according to actual needs. This embodiment of the application does not impose any restrictions.

[0163] It is understandable that when the ratio corresponding to the characters in line L of the first paragraph is greater than the first threshold, the characters in line L are likely to have first-line indentation. Therefore, it can be determined that line L is the beginning of a paragraph, and thus line L and the line preceding line L belong to different paragraphs. This process can be repeated to check if the ratio corresponding to each line of characters in the first paragraph is greater than the second threshold, in order to determine whether each line of characters and the line preceding it should be divided into different paragraphs.

[0164] In this way, the electronic device can use the difference between the coordinate information of the first character in each line and the coordinate information of the first character in the previous line, as well as the average width of the characters in each line, to perform secondary paragraph division on the text in the first image, thereby improving the accuracy of paragraph division.

[0165] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 10 As shown, step 203 above can be implemented through step 203a below.

[0166] Step 203a: The electronic device plays the text in the first image sequentially, according to at least one paragraph.

[0167] In this embodiment of the application, the electronic device can play the text in the first image based on at least one paragraph.

[0168] In this embodiment, the electronic device can play the text segments in the first image recognized in a top-to-bottom order, or it can play the segments in the first image recognized in a random order. This embodiment does not impose any specific limitations.

[0169] Optionally, in the embodiments of this application, step 203 above can be implemented by step 203b below.

[0170] Step 203b: The electronic device plays the text in the first image based on the paragraph selected by the user from at least one paragraph.

[0171] In this embodiment, the user can select the paragraph to be played by touch input or voice input. Then, the electronic device can play the text in the paragraph selected by the user in the order of the selected paragraphs.

[0172] Optionally, in this embodiment, the touch input may include, but is not limited to, the user using a touch device such as a finger or stylus to input data onto the electronic device, such as a specific gesture or other feasible input. The specific input can be determined based on actual usage needs, and this embodiment does not impose any limitations.

[0173] Optionally, in the embodiments of this application, the aforementioned specific gesture can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-tap gesture.

[0174] In this way, electronic devices can play the recognized text content in a certain order, so that users can understand the content and structure of the text in the first image through hearing, thus improving the flexibility of electronic devices in playing text content.

[0175] The implementation process of the image recognition method provided in this application embodiment will be illustrated below through specific implementation methods.

[0176] This implementation method, based on image processing, draws the text region as color blocks onto a blank image, then uses image aggregation to calculate the text paragraph region, ultimately obtaining the text paragraph range. Subsequently, according to the rules of standard paragraphs, natural paragraphs are identified and further segmented, effectively solving the problem that OCR technology can only read text line by line. Figure 11 As shown, the specific steps are as follows: steps 21 to 36.

[0177] Step 21: The electronic device recognizes the text through the camera and prompts the user.

[0178] It should be noted that when an electronic device detects text, it will summarize the text content and read the summary aloud to inform the user of the general content of the text.

[0179] Step 22: The electronic device displays a text snapshot button and prompts the user to click to get the detailed content of the text.

[0180] Step 23: The electronic device detects whether the user clicks the text snapshot button.

[0181] If yes, proceed to step 24 below; otherwise, repeat step 22 above.

[0182] It is understandable that if the user clicks the button, the electronic device will take a picture of the content detected by the current camera, obtain a first image, and then transmit the first image to the corresponding OCR module in the electronic device.

[0183] Step 24: The electronic device recognizes the first image and obtains the coordinate information of the first and last characters of each line in the first image.

[0184] Step 25: The electronic device uses the first color block to mark the area where each line of characters is located based on the obtained coordinate information.

[0185] Step 26: The electronic device calculates the average line height of the characters in the first image.

[0186] Step 27: The electronic device determines whether there is a character line in the first image whose ratio of line height to average line height is greater than a first threshold.

[0187] If yes, proceed to step 28 below; if no, proceed to step 29 below.

[0188] Step 28: The electronic device calculates the average line height of the remaining character lines, excluding character lines whose line height to average line height ratio is greater than the first threshold.

[0189] Step 29: The electronic device adjusts the height of the first color block according to the average line height of the characters to obtain the second color block corresponding to each first color block.

[0190] Step 30: The electronic device identifies the overlapping second color blocks as a third color block and marks the outline of each third color block based on the outline coordinate information of each third color block.

[0191] Step 31: The electronic device obtains the coordinate information of the four vertices corresponding to the four boundaries of the outline of each third color block.

[0192] Step 32: Based on the coordinate information of the four vertices, the electronic device marks the border corresponding to each third color block, and divides the text in the first image into M paragraphs.

[0193] Step 33: The electronic device calculates the ratio between the difference between the coordinates of the first character in each row and the coordinates of the first character in the previous row within the border corresponding to each third color block, and the average width of each row of characters.

[0194] Step 34: The electronic device determines whether there is a ratio greater than the second threshold within the border corresponding to each third color block and at least one line of characters.

[0195] If yes, proceed to step 35 below; if no, proceed to step 36 below.

[0196] Step 35: The electronic device divides the character lines whose ratios are greater than the second threshold into different paragraphs with the line above them, so as to divide M paragraphs into N paragraphs.

[0197] Step 36: The electronic device plays the text in the first image based on N paragraphs.

[0198] In this embodiment, the electronic device can use image processing to draw the text region as color blocks onto a blank image, and then use image aggregation to calculate the text paragraph region, ultimately obtaining the text paragraph range. Subsequently, the electronic device identifies natural paragraphs according to the rules of standard paragraphs, performs secondary segmentation, obtains better recognition results, realizes the aggregation of natural paragraphs of text, and effectively solves the problem that OCR recognition can only read text line by line.

[0199] Each of the above-described method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application does not impose any restrictions on this.

[0200] The image recognition method provided in this application can be executed by an image recognition device. This application uses an image recognition device executing the image recognition method as an example to illustrate the image recognition device provided in this application.

[0201] Figure 12 A schematic diagram of a possible structure of the image recognition device involved in some embodiments of this application is shown. For example... Figure 12 As shown, the image recognition device 70 may include a processing module 71 and a playback module 72.

[0202] The processing module 71 is used to identify the first image and obtain the coordinate information of each line of characters in the first image; the processing module 71 is also used to divide the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image; the playback module 72 is used to play the text in the first image based on the at least one paragraph obtained by the processing module 71.

[0203] In one possible implementation, the coordinate information includes the coordinate information of the first and last characters in each line of characters; the processing module 71 is specifically used to mark the area where each line of characters is located using a first color block based on the coordinate information of the first and last characters in each line of characters, and to adjust the height of the first color block corresponding to each line of characters based on the average line height of the characters in the first image to obtain at least one second color block; and to divide the characters contained in the overlapping second color blocks into the same paragraph to obtain at least one paragraph.

[0204] In one possible implementation, the processing module 71 is specifically configured to: determine the overlapping second color blocks among the at least one second color block as a third color block, thereby obtaining at least one third color block; and, based on the contour coordinate information of each third color block, obtain the four boundary information of the contour of each third color block; and, based on the four boundary information of the contour of each third color block, obtain the coordinate information of the four vertices corresponding to the four boundary information; and, based on the coordinate information of the four vertices, mark the border corresponding to each third color block; and, divide the characters contained in each border into the same paragraph, thereby obtaining at least one paragraph.

[0205] In one possible implementation, the processing module 71 is further configured to calculate the line height of each line of characters based on the coordinate information of any character in each line, wherein the coordinate information of any character includes the upper boundary coordinate information and the lower boundary coordinate information of any character; and to calculate a first average line height based on the line height of each line of characters; and,

[0206] If the ratio of the row height of all rows of characters in the first image to the first average row height is less than or equal to the first threshold, then the first average row height is determined as the average row height of the characters in the first image.

[0207] If the ratio of the height of at least one line of characters to the first average line height is greater than the first threshold, then the average line height of the characters in the first image is calculated based on the coordinate information of the remaining lines of characters in the first image excluding the at least one line of characters.

[0208] In one possible implementation, the above-mentioned at least one paragraph is N paragraphs, where N is an integer greater than 1; the above-mentioned processing module 71 is specifically used to divide the characters in the first image into M paragraphs based on coordinate information and the average line height of the characters in the first image, where M is a positive integer; and to divide the characters in the M paragraphs into N paragraphs based on the coordinate information and average width of each line of characters in the M paragraphs, where N is greater than M.

[0209] In one possible implementation, the processing module 71 is specifically used to calculate the difference between the coordinate information of the first character of each line in the first paragraph and the coordinate information of the first character of the previous line, and to calculate the ratio between this difference and the average width of each line of characters in the first paragraph; wherein the first paragraph is any one of the M paragraphs; and,

[0210] If the ratio corresponding to the Lth line character in the first paragraph is greater than the second threshold, then the Lth line character and the line character above the Lth line character are divided into different paragraphs to obtain N paragraphs, where L is a positive integer.

[0211] In one possible implementation, the playback module 72 is specifically used to play the text in the first image sequentially according to the order of at least one paragraph;

[0212] or,

[0213] Based on the paragraph selected by the user from at least one of the above paragraphs, the text in the first image is played.

[0214] This application provides an image recognition device. When recognizing an image, the image recognition device can determine the position of each character in a line based on the coordinate information of each character. Then, based on the line height of the character at each position, it can determine whether there is a paragraph between the lines, thereby dividing the characters into paragraphs. This allows characters within the same paragraph to be grouped together, and thus, during voice broadcasting, it can continuously broadcast the text content of a natural paragraph, enabling users to accurately understand the text content in the image through hearing. This improves the flexibility of the image recognition device in recognizing characters in the image and provides a better effect for broadcasting text to users.

[0215] The image recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0216] The image recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0217] The image recognition device provided in this application embodiment can implement all the processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0218] Optionally, such as Figure 13 As shown, this application embodiment also provides an electronic device 1000, including a processor 1001 and a memory 1002. The memory 1002 stores a program or instructions that can run on the processor 1001. When the program or instructions are executed by the processor 1001, they implement the various steps of the above-described image recognition method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0219] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0220] Figure 14 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0221] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0222] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 14 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0223] The processor 110 is used to recognize the first image and obtain the coordinate information of each line of characters in the first image; the processor 110 is also used to divide the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image; the audio output unit 103 is used to play the text in the first image based on the at least one paragraph obtained by the processor 110.

[0224] In one possible implementation, the coordinate information includes the coordinate information of the first and last characters in each line of characters; the processor 110 is specifically used to mark the area where each line of characters is located using a first color block based on the coordinate information of the first and last characters in each line of characters; and to adjust the height of the first color block corresponding to each line of characters based on the average line height of the characters in the first image to obtain at least one second color block; and to divide the characters contained in the overlapping second color blocks into the same paragraph to obtain at least one paragraph.

[0225] In one possible implementation, the processor 110 is specifically configured to: determine the overlapping second color blocks among the at least one second color block as a third color block, thereby obtaining at least one third color block; and, based on the contour coordinate information of each third color block, obtain the four boundary information of the contour of each third color block; and, based on the four boundary information of the contour of each third color block, obtain the coordinate information of the four vertices corresponding to the four boundary information; and, based on the coordinate information of the four vertices, mark the border corresponding to each third color block; and, divide the characters contained in each border into the same paragraph, thereby obtaining at least one paragraph.

[0226] In one possible implementation, the processor 110 is further configured to calculate the line height of each line of characters based on the coordinate information of any character in each line, wherein the coordinate information of any character includes the upper boundary coordinate information and the lower boundary coordinate information of any character; and to calculate a first average line height based on the line height of each line of characters; and,

[0227] If the ratio of the row height of all rows of characters in the first image to the first average row height is less than or equal to the first threshold, then the first average row height is determined as the average row height of the characters in the first image.

[0228] If the ratio of the height of at least one line of characters to the first average line height is greater than the first threshold, then the average line height of the characters in the first image is calculated based on the coordinate information of the remaining lines of characters in the first image excluding the at least one line of characters.

[0229] In one possible implementation, the above-mentioned at least one paragraph is N paragraphs, where N is an integer greater than 1; the processor 110 is specifically used to divide the characters in the first image into M paragraphs based on coordinate information and the average line height of the characters in the first image, where M is a positive integer; and to divide the characters in the M paragraphs into N paragraphs based on the coordinate information and average width of each line of characters in the M paragraphs, where N is greater than M.

[0230] In one possible implementation, the processor 110 is specifically configured to calculate the difference between the coordinate information of the first character of each line in the first paragraph and the coordinate information of the first character of the previous line, and to calculate the ratio between this difference and the average width of each line of characters in the first paragraph; wherein the first paragraph is any one of M paragraphs; and,

[0231] If the ratio corresponding to the Lth line character in the first paragraph is greater than the second threshold, then the Lth line character and the line character above the Lth line character are divided into different paragraphs to obtain N paragraphs, where L is a positive integer.

[0232] In one possible implementation, the audio output unit 103 is specifically used to play the text in the first image sequentially according to the order of at least one paragraph;

[0233] or,

[0234] Based on the paragraph selected by the user from at least one of the above paragraphs, the text in the first image is played.

[0235] This application provides an electronic device that, when recognizing an image, can determine the position of each character based on the coordinate information of each line of characters. Then, based on the line height of each character, it can determine whether the lines cross paragraphs, thereby dividing the characters into paragraphs. This allows characters within the same paragraph to be grouped together, enabling continuous playback of a natural paragraph of text during voice broadcasting. This allows users to accurately understand the text content in the image through hearing, thus improving the flexibility of the electronic device in recognizing characters in the image and providing a better effect for playing text for the user.

[0236] The electronic device provided in this application embodiment can implement all the processes implemented in the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here. The beneficial effects of the various implementation methods in this embodiment can be found in the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, it will not be described again here.

[0237] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0238] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0239] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0240] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image recognition method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0241] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0242] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image recognition method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0243] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0244] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the image recognition method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0245] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0246] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0247] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image recognition method, characterized in that, include: The first image is identified to obtain the coordinate information of each line of characters in the first image; Based on the coordinate information and the average line height of the characters in the first image, the characters in the first image are divided into at least one paragraph; Based on the at least one paragraph, the text in the first image is played.

2. The method according to claim 1, characterized in that, The coordinate information includes the coordinate information of the first and last characters in each line of characters; The step of dividing the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image includes: Based on the coordinate information of the first and last characters in each line of characters, the area where each line of characters is located is marked with a first color block; Based on the average row height of the characters in the first image, the height of the first color block corresponding to each row of characters is adjusted to obtain at least one second color block; The characters contained in the overlapping second color blocks are divided into the same paragraph to obtain the at least one paragraph.

3. The method according to claim 2, characterized in that, The step of dividing the characters contained in overlapping second color blocks into the same paragraph to obtain the at least one paragraph includes: The overlapping second color blocks among the at least one second color blocks are identified as a third color block, thus obtaining at least one third color block; Based on the contour coordinate information of each third color block, obtain the four boundary information of the contour of each third color block; Based on the four boundary information of the outline of each third color block, obtain the coordinate information of the four vertices corresponding to the four boundary information; Based on the coordinate information of the four vertices, mark the border corresponding to each third color block; The characters contained in each of the borders are divided into the same paragraph to obtain at least one paragraph.

4. The method according to claim 2, characterized in that, The method further includes: Based on the coordinate information of any character in each line of characters, the line height of each line of characters is calculated, and the coordinate information of any character includes the upper boundary coordinate information and the lower boundary coordinate information of any character. Calculate the first average row height based on the row height of each row of characters; If the ratio of the row height of all rows of characters in the first image to the first average row height is less than or equal to the first threshold, then the first average row height is determined as the average row height of the characters in the first image. If the ratio of the height of at least one line of characters to the first average line height is greater than the first threshold, then the average line height of the characters in the first image is calculated based on the coordinate information of the remaining lines of characters in the first image excluding the at least one line of characters.

5. The method according to any one of claims 1 to 4, characterized in that, The at least one paragraph consists of N paragraphs, where N is an integer greater than 1; The step of dividing the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image includes: Based on the coordinate information and the average line height of the characters in the first image, the characters in the first image are divided into M segments, where M is a positive integer; Based on the coordinate information and average width of each line of characters in the M paragraphs, the characters in the M paragraphs are divided into N paragraphs, where N is greater than M.

6. The method according to claim 5, characterized in that, The characters in the M paragraphs are divided based on the coordinate information and average width of each line of characters in the M paragraphs to obtain the N paragraphs, including: Calculate the difference between the coordinate information of the first character of each line in the first paragraph and the coordinate information of the first character of the previous line, where the first paragraph is any one of the M paragraphs; Calculate the ratio between the difference and the average width of each line of characters in the first paragraph; If the ratio corresponding to the Lth line character in the first paragraph is greater than the second threshold, then the Lth line character and the previous line character of the Lth line character are divided into different paragraphs to obtain the N paragraphs, where L is a positive integer.

7. An image recognition device, characterized in that, The device includes: a processing module and a playback module; The processing module is used to recognize the first image and obtain the coordinate information of each line of characters in the first image; The processing module is further configured to divide the characters in the first image into at least one paragraph based on the coordinate information and the average line height of the characters in the first image; The playback module is used to play the text in the first image based on the at least one paragraph obtained by the processing module.

8. The apparatus according to claim 7, characterized in that, The coordinate information includes the coordinate information of the first and last characters in each line of characters; The processing module is specifically used to mark the area where each line of characters is located using a first color block based on the coordinate information of the first and last characters in each line of characters; Furthermore, based on the average row height of the characters in the first image, the height of the first color block corresponding to each row of characters is adjusted to obtain at least one second color block; And, the characters contained in the overlapping second color blocks in the at least one second color block are divided into the same paragraph to obtain the at least one paragraph.

9. The apparatus according to claim 8, characterized in that, The processing module is specifically configured to determine the overlapping second color blocks among the at least one second color block as a third color block, thereby obtaining at least one third color block; and, based on the contour coordinate information of each third color block, obtain the four boundary information of the contour of each third color block; and, based on the four boundary information of the contour of each third color block, obtain the coordinate information of the four vertices corresponding to the four boundary information. Furthermore, based on the coordinate information of the four vertices, the border corresponding to each third color block is marked; and the characters contained in each border are divided into the same paragraph to obtain the at least one paragraph.

10. The apparatus according to claim 8, characterized in that, The processing module is further configured to calculate the line height of each line of characters based on the coordinate information of any character in each line of characters, wherein the coordinate information of any character includes the upper boundary coordinate information and the lower boundary coordinate information of any character. And, based on the line height of each line of characters, calculate the first average line height; The processing module is further configured to determine the first average line height as the average line height of the characters in the first image if the ratio of the line height of all lines of characters in the first image to the first average line height is less than or equal to the first threshold; or, if the ratio of the line height of at least one line of characters to the first average line height is greater than the first threshold, calculate the average line height of the characters in the first image based on the coordinate information of the remaining lines of characters in the first image other than the at least one line of characters.

11. The apparatus according to any one of claims 7 to 10, characterized in that, The at least one paragraph consists of N paragraphs, where N is an integer greater than 1; The processing module is specifically used to divide the characters in the first image into M segments based on the coordinate information and the average line height of the characters in the first image, where M is a positive integer; Furthermore, based on the coordinate information and average width of each line of characters in the M paragraphs, the characters in the M paragraphs are divided to obtain the N paragraphs, where N is greater than M.

12. The apparatus according to claim 11, wherein the processing module is specifically configured to calculate the difference between the coordinate information of the first character of each line in the first paragraph and the coordinate information of the first character of the previous line, and to calculate the ratio between the difference and the average width of the characters in each line of the first paragraph; wherein, The first paragraph is any one of the M paragraphs; and if the ratio corresponding to the Lth row of characters in the first paragraph is greater than the second threshold, then the Lth row of characters and the previous row of characters in the Lth row of characters are divided into different paragraphs to obtain the N paragraphs, where L is a positive integer.

13. The apparatus according to claim 7, characterized in that, The playback module is specifically used to play the text in the first image sequentially according to the order of the at least one paragraph; or, based on the paragraph selected by the user from the at least one paragraph, to play the text in the first image.