Information Recognition Method, Device, Electronic Device and Storage Medium

By acquiring and stitching image features, text content features and position features, adjusting the output position of multiple lines of text, the problem of incoherent identification results in the prior art is solved, and more accurate text recognition results are achieved.

CN113920293BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111210987.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-06-10
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify the output location of multiple lines of text in an image, resulting in the recognition results not meeting the needs of users.

Method used

By obtaining the image features, text content features and position features of each line of text in the image to be identified, splicing them into multimodal features, and inputting them into the trained preset model, outputting the pointer position corresponding to each line of text to adjust the output position.

Benefits of technology

Improve the accuracy of output position of multiple lines of text, so that the recognition results maintain semantic coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920293B_ABST
    Figure CN113920293B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information recognition method, apparatus, electronic device, and storage medium. The method includes: obtaining an image to be recognized; obtaining the image features, text content features, and position features of the text content features in the image to be recognized for each line of text; splicing the image features, text content features, and position features of each line of text to obtain multi-modal features of multiple lines of text; inputting the multi-modal features of multiple lines of text into a trained preset model, and outputting the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content features in the image to be recognized for each line of text are combined to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition result can maintain semantic coherence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information technology, and in particular, to an information recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of technology, OCR (Optical Character Recognition) technology has also made great progress, and users can easily recognize the text on images through relevant applications. However, for pictures with rich information in the text or background, such as posters or certain video frames, the results recognized by existing technologies often cannot meet the needs of users. Summary of the Invention

[0003] To overcome the problems in the related art, the present disclosure provides an information recognition method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, an information recognition method is provided, including:

[0005] Obtain an image to be recognized, where the image to be recognized includes multiple lines of text;

[0006] Respectively obtain the image feature, text content feature, and position feature of the text content in the image to be recognized for each line of text;

[0007] Concatenate the image feature, text content feature, and position feature of each line of text to obtain a multi-modal feature of the multiple lines of text;

[0008] Input the multi-modal feature of the multiple lines of text into a trained preset model, and output the pointer position corresponding to each line of text, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text.

[0009] Optionally, the training process of the preset model includes:

[0010] Obtain a training sample image, where the training sample image contains multiple lines of text;

[0011] Extract the sample feature of the training sample image, where the sample feature includes the image feature, text content feature, and position feature of the multiple lines of text in the training sample image;

[0012] Input the sample features into a preset model to obtain the sample pointer positions corresponding to each line of text in the training sample image. Calculate the loss value between the sample pointer positions and the labeled pointer positions corresponding to each line of text in the pre-labeled training sample image through an objective loss function. When the loss value is less than a threshold, obtain the trained preset model.

[0013] Optionally, the method further includes:

[0014] Based on the pointer positions corresponding to each line of text in the image to be recognized, sort the multiple lines of text in the image to be recognized to obtain a sorting result;

[0015] Display the multiple lines of text included in the image to be recognized according to the sorting result.

[0016] Optionally, the obtaining of the image features of each line of text includes:

[0017] Extract the image features of each line of text through a convolutional neural network. The image features include one or a combination of text size features, text color features, texture features, and background features.

[0018] Optionally, the image features, text content features, and the position features are respectively 256-dimensional feature vectors, and the multi-modal features are 256*3-dimensional feature vectors obtained by splicing the image features, text content features, and position features.

[0019] According to a second aspect of the embodiments of the present disclosure, an information recognition device is provided, including:

[0020] An image acquisition module configured to acquire an image to be recognized, where the image to be recognized includes multiple lines of text;

[0021] A feature acquisition module configured to respectively acquire the image features, text content features, and the position features of the text content in the image to be recognized;

[0022] A feature splicing module configured to splice the image features, text content features, and position features of each line of text to obtain the multi-modal features of the multiple lines of text;

[0023] A position acquisition module configured to input the multi-modal features of the multiple lines of text into the trained preset model and output the pointer positions corresponding to each line of text, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text.

[0024] Optionally, it further includes a feature training module, and the feature training module is specifically configured to execute:

[0025] Obtain training sample images, where the training sample images contain multiple lines of text;

[0026] Extract the sample features of the training sample images, where the sample features include the image features, text content features of the multiple lines of text in the training sample images, and the position features of the text content features in the corresponding sample images;

[0027] Input the sample features into a preset model to obtain the sample pointer positions corresponding to each line of text in the training sample images, calculate the loss value between the sample pointer positions and the labeled pointer positions corresponding to each line of text in the pre-labeled training sample images through a target loss function, and when the loss value is less than a threshold, obtain the trained preset model.

[0028] Optionally, the device further includes:

[0029] A text sorting module configured to perform sorting on the multiple lines of text in the image to be recognized based on the pointer positions corresponding to each line of text in the image to be recognized, and obtain a sorting result;

[0030] A text display module configured to perform displaying the multiple lines of text included in the image to be recognized according to the sorting result.

[0031] Optionally, the feature acquisition module is specifically configured to perform:

[0032] Extract the image features of each line of text through a convolutional neural network, where the image features include one or a combination of several of: text size features, text color features, texture features, and background features.

[0033] Optionally, the image features, text content features, and the position features are respectively 256-dimensional feature vectors, and the multi-modal features are 256*3-dimensional feature vectors obtained by splicing the image features, text content features, and the position features.

[0034] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0035] A processor;

[0036] A memory for storing processor-executable instructions;

[0037] Wherein, the processor is configured to execute the information recognition method described in the first aspect.

[0038] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the steps of the information recognition method described in the first aspect.

[0039] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product. The computer program product includes computer instructions. When the computer instructions run on an electronic device, the electronic device is enabled to execute the steps of the information recognition method described in the first aspect.

[0040] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0041] For the technical solution provided by the embodiment of the present disclosure, an image to be recognized is obtained; image features, text content features, and position features of the text content features in the image to be recognized are respectively obtained for each row of text in the image to be recognized; the image features, text content features, and position features of each row of text are spliced to obtain multi-modal features of multiple rows of text; the multi-modal features of multiple rows of text are input into a trained preset model, and pointer positions respectively corresponding to each row of text are output, where the pointer position corresponding to each row of text is used to represent the output position corresponding to that row of text. It can be seen that through the technical solution provided by the embodiment of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content features in the image to be recognized for each row of text will be combined simultaneously to adjust the output positions of multiple rows of text, so that the output positions of multiple rows of text are more accurate, and thus the obtained text recognition result can maintain semantic coherence.

[0042] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0044] Figure 1 is a flowchart of an information recognition method shown according to an exemplary embodiment;

[0045] Figure 2 is another flowchart of an information recognition method shown according to an exemplary embodiment;

[0046] Figure 3 is a block diagram of an information recognition device shown according to an exemplary embodiment;

[0047] Figure 4It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0048] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0049] Figure 1 It is a flowchart of an information recognition method shown according to an exemplary embodiment. As Figure 1 shown, this method is used in a terminal and may include the following steps:

[0050] In step S110, an image to be recognized is acquired.

[0051] Among them, the image to be recognized includes multiple lines of text.

[0052] Specifically, the OCR technology realizes detecting and recognizing the text in the image into text. In most cases, OCR recognizes text line by line. In practical applications, an image usually includes multiple lines of text. For example, in the short video application scenario, a frame of an image usually includes multiple lines of text, and different text lines represent different meanings.

[0053] In step S120, the image feature, the text content feature, and the position feature of the text content in the image to be recognized are respectively acquired for each line of text.

[0054] Specifically, since the image to be recognized usually contains rich information, in addition to text information, there will also be background images, etc. In addition, the text on the image to be recognized usually has fonts of different sizes. For the same-sized text, even if it is displayed on different lines, it often represents semantic coherence. In addition, text with the same background and texture is usually semantically coherent. Therefore, when recognizing multiple lines of text in the image to be recognized, it is necessary to acquire the image feature, the text content feature, and the position feature of the text content in the image to be recognized for each line of text, so that in the subsequent steps, the output positions of the multiple lines of text in the image to be recognized can be accurately determined.

[0055] In order to accurately obtain the image feature of each line of text, in one implementation manner, acquiring the image feature of each line of text may include the following steps:

[0056] Extract the image features of each line of text through a convolutional neural network. The image features include one or a combination of the following: text size feature, text color feature, texture feature, and background feature.

[0057] In this embodiment, the image features of the text region in the image to be recognized can be extracted through a convolutional neural network CNN. Among them, the image features include one or a combination of the following: text size feature, text color feature, texture feature, and background feature. It can be seen that the image features of each line of text are not just single features, but can be multiple features such as text size feature, text color feature, texture feature, and background feature, so that the obtained image features are more accurate.

[0058] Moreover, when obtaining the text content features of each line of text, multiple lines of text in the image to be recognized can be input into the trained Bert model, and the trained Bert model is used to recognize multiple lines of text in the image to be recognized to obtain the text content features of each line of text. And the position features of the text content features in the image to be recognized can be extracted by using multi-level cascading of the coordinates of the text lines.

[0059] In step S130, the image features, text content features, and position features of each line of text are concatenated to obtain the multi-modal features of multiple lines of text.

[0060] Specifically, in order to accurately determine the output position of each line of text in the subsequent steps, the image features, text content features, and position features of each line of text can be concatenated to obtain the multi-modal features of each line of text. Since the multi-modal features of each line of text are concatenated based on these three features: image features, text content features, and position features, in the subsequent steps, the output positions of multiple lines of text can be determined more accurately based on the multi-modal features of multiple lines of text.

[0061] In one embodiment, the image features, text content features, and position features are respectively 256-dimensional feature vectors, and the multi-modal feature is a 256*3-dimensional feature vector obtained by concatenating the image features, text content features, and position features.

[0062] In this embodiment, the image features, text content features, and position features of each line of text can be respectively represented as 256-dimensional feature vectors. By concatenating these three 256-dimensional feature vectors, a 256*3-dimensional multi-modal feature vector can be obtained, that is, the multi-modal feature can be a 768-dimensional feature vector.

[0063] As can be seen from the above description, the image features, text content features, and position features are all feature vectors with relatively high dimensions. Therefore, the accuracy of the image features, text content features, and position features is relatively high. Then, the image features, text content features, and position features are concatenated, and the accuracy of the resulting multi-modal features is also relatively high. That is, the multi-modal features of each line of text can accurately represent the features of that line of text, which is conducive to accurately determining the output position of that line of text based on the multi-modal features of that line of text.

[0064] In step S140, the multi-modal features of multiple lines of text are input into the trained preset model, and the pointer positions corresponding to each line of text are output.

[0065] Among them, the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text.

[0066] In the embodiments provided by the present disclosure, the preset model can be a Pointer Network. By training the Pointer Network with multiple labeled samples, the trained preset model can be obtained. By inputting the multi-modal features obtained by concatenating the image features, text content features, and position features into the trained preset model, the pointer positions corresponding to each line of text can be obtained, and the pointer positions corresponding to each line of text are used to represent the output positions corresponding to that line of text.

[0067] Exemplarily, the image to be recognized includes 3 lines of text, and the pointer positions corresponding to each line of text in the output result are 2, 1, and 1 respectively, which respectively represent the output line numbers corresponding to these 3 lines of text. That is, the output position of the original first line of text in the image to be recognized is the second line, and the output positions of the original second line of text and the original third line of text are both the first line. Of course, this is just a simple example here. The purpose of the embodiments of the present disclosure is to output the text with the same text size, color, and other features on the same line. By adjusting the output positions of multiple lines of text, the semantic coherence of the output text can be maintained.

[0068] The technical solution provided by the embodiments of the present disclosure obtains an image to be recognized; respectively obtains the image features, text content features, and position features of the text content in the image to be recognized for each line of text; splices the image features, text content features, and position features of each line of text to obtain the multi-modal features of multiple lines of text; inputs the multi-modal features of multiple lines of text into a trained preset model, and outputs the pointer positions corresponding to each line of text, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that through the technical solution provided by the embodiments of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content in the image to be recognized for each line of text will be combined simultaneously to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition results can maintain semantic coherence.

[0069] Combined with the above embodiments, in another embodiment provided by the present disclosure, as Figure 2 shown, the training process of the above preset model may include the following steps:

[0070] Step S210, obtain training sample images.

[0071] Among them, the training sample images contain multiple lines of text.

[0072] Specifically, in order to enable the trained model to accurately output the output positions corresponding to multiple lines of text in an image, when obtaining training sample images, it is necessary to obtain a large number of images containing multiple lines of text, and the sample images contain corresponding image backgrounds and textures, and text lines with different sizes of text.

[0073] Moreover, after obtaining the training sample images, the output positions of each line of text can be pre-annotated. Specifically, the multiple lines of text included in each training sample image are known, so the output positions of multiple lines of text can be annotated, and the text lines output through the annotated output positions have semantic coherence.

[0074] Step S220, extract the sample features of the training sample images.

[0075] Among them, the sample features include the image features, text content features, and position features of the multiple lines of text in the training sample images.

[0076] Specifically, after obtaining the training sample images, the sample features of the training sample images can be extracted. Specifically, three network models can be used to extract them respectively. That is, the recognition results of the text lines are extracted through the trained Bert model to obtain a 256-dimensional feature vector; the trained convolutional neural network CNN is used to extract the image features in the text region to obtain a 256-dimensional feature vector; the text line coordinates are extracted by multi-stage cascaded FC to obtain a 256-dimensional feature vector.

[0077] Step S230: Input the sample features into the preset model to obtain the sample pointer positions corresponding to each line of text in the training sample image. Calculate the loss value between the sample pointer position and the labeled pointer position corresponding to each line of text in the pre-labeled training sample image through the target loss function. When the loss value is less than the threshold, the trained preset model is obtained.

[0078] Specifically, after obtaining the sample features of the training sample image, the sample features of the training sample image can be input into the preset model to train the preset model. What is output from the preset model is the sample pointer position corresponding to each line of text. Since the labeled pointer position corresponding to each line of text in the training sample image is pre-labeled, and the labeled pointer position corresponding to each line of text is used to represent the output position when the semantics is coherent, that is to say, the labeled pointer position is the true value. Therefore, when training the preset model, the loss value between the sample pointer position and the labeled pointer position corresponding to each line of text in the pre-labeled training sample image is calculated through the target loss function. When the loss value is less than the threshold, it means that the sample pointer position output from the preset model is close to the pre-labeled labeled pointer position. That is to say, the accuracy of the preset model is relatively high. At this time, the trained preset model is obtained.

[0079] As can be seen from the above description, when training the preset model, the sample features of the training sample image are also multi-modal features, and the dimension of the multi-modal features is relatively high, that is, the multi-modal features are relatively accurate. And the sample pointer position output from the trained preset model is close to the pre-labeled labeled pointer position. That is to say, the accuracy of the trained preset model is relatively high. Furthermore, after inputting the image to be recognized into the trained preset model, the accuracy of the output positions of multiple lines of text in the image to be recognized is relatively high, and thus the obtained text recognition result can maintain semantic coherence.

[0080] Combined with the above embodiments, in another embodiment provided by the present disclosure, the information recognition method may further include the following steps:

[0081] Step a1: Sort the multiple lines of text in the image to be recognized based on the pointer positions corresponding to each line of text in the image to be recognized to obtain a sorting result.

[0082] Step a2: Display the multiple lines of text included in the image to be recognized according to the sorting result.

[0083] Specifically, since the pointer position corresponding to each line of text can be used to represent the output position corresponding to that line of text, after obtaining the pointer positions corresponding to the multiple lines of text, that is, obtaining the output positions corresponding to the multiple lines of text. Therefore, after obtaining the pointer positions corresponding to each line of text respectively, the multiple lines of text can be sorted according to the pointer positions corresponding to each line of text respectively to obtain the sorting result, and the multiple lines of text are displayed according to the sorting result.

[0084] For example, the image to be recognized includes 3 lines of text, and the pointer positions corresponding to each line of text in the output result are 3, 1, and 2 respectively, which represent the output positions corresponding to these 3 lines of text, that is, the output position of the original first line of text in the image to be recognized is the third line, the output position of the original second line of text is the first line, and the output position of the original third line of text is the second line. These three lines of text are displayed after being sorted according to the pointer positions, that is, the second line of text, the third line of text, and the first line of text are displayed in sequence.

[0085] It can be seen that through the technical solution provided by the embodiments of the present disclosure, the output positions of multiple lines of text can be adjusted, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition result can maintain semantic coherence.

[0086] Figure 3 It is a block diagram of an information recognition device shown according to an exemplary embodiment. Refer to Figure 3 , the device includes an image acquisition module 310 configured to acquire an image to be recognized, where the image to be recognized includes multiple lines of text;

[0087] A feature acquisition module 320 configured to respectively acquire the image feature, the text content feature, and the position feature of the text content feature in the image to be recognized for each line of text;

[0088] A feature splicing module 330 configured to splice the image feature, the text content feature, and the position feature of each line of text to obtain the multi-modal feature of the multiple lines of text;

[0089] A position acquisition module 340 configured to input the multi-modal feature of the multiple lines of text into a trained preset model and output the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text.

[0090] Optionally, it further includes a feature training module, and the feature training module is specifically configured to perform:

[0091] Acquire a training sample image, where the training sample image contains multiple lines of text;

[0092] Extract the sample features of the training sample image, where the sample features include the image features of multiple lines of text in the training sample image, the text content features, and the position features of the text content features in the corresponding sample image;

[0093] Input the sample features into a preset model to obtain the sample pointer positions corresponding to each line of text in the training sample image. Calculate the loss value between the sample pointer positions and the labeled pointer positions corresponding to each line of text in the pre-labeled training sample image through an objective loss function. When the loss value is less than the threshold, obtain the trained preset model.

[0094] Optionally, the device further includes:

[0095] A text sorting module, configured to perform sorting on the multiple lines of text in the image to be recognized based on the pointer positions corresponding to each line of text in the image to be recognized, and obtain a sorting result;

[0096] A text display module, configured to perform displaying the multiple lines of text included in the image to be recognized according to the sorting result.

[0097] Optionally, the feature acquisition module is specifically configured to perform:

[0098] Extract the image features of each line of text through a convolutional neural network, where the image features include one or several combinations of text size features, text color features, texture features, and background features.

[0099] Optionally, the image features, text content features, and the position features are respectively 256-dimensional feature vectors, and the multi-modal features are 256*3-dimensional feature vectors obtained by splicing the image features, text content features, and the position features.

[0100] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0101] The technical solution provided by the embodiments of the present disclosure obtains an image to be recognized; respectively obtains the image features, text content features, and position features of the text content in the image to be recognized for each line of text; splices the image features, text content features, and position features of each line of text to obtain the multi-modal features of multiple lines of text; inputs the multi-modal features of multiple lines of text into a trained preset model, and outputs the pointer positions corresponding to each line of text, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that through the technical solution provided by the embodiments of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content in the image to be recognized for each line of text will be combined simultaneously to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition results can maintain semantic coherence.

[0102] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0103] A processor;

[0104] A memory for storing processor-executable instructions;

[0105] Wherein, the processor is configured to execute the information recognition method described in the first aspect.

[0106] Figure 4 FIG. is a block diagram of an information recognition device 800 shown according to an exemplary embodiment. For example, the device 800 is an electronic device, specifically, it can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0107] Referring to Figure 4 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0108] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0109] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0110] The power supply component 806 provides power to various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0111] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0112] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0113] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0114] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, the sensor assembly 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0115] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0116] In an exemplary embodiment, the device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above... methods.

[0117] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by the processor 820 of the device 800 to complete the above methods. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0118] The technical solution provided by the embodiments of the present disclosure obtains an image to be recognized; respectively obtains the image features, text content features, and position features of the text content features in the image to be recognized for each line of text; splices the image features, text content features, and position features of each line of text to obtain multi-modal features of multiple lines of text; inputs the multi-modal features of multiple lines of text into a trained preset model, and outputs the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that through the technical solution provided by the embodiments of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content features in the image to be recognized for each line of text will be combined simultaneously to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition result can maintain semantic coherence.

[0119] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, which when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the steps of the information recognition method described in the first aspect.

[0120] The technical solution provided by the embodiments of the present disclosure obtains an image to be recognized; respectively obtains the image features, text content features, and position features of the text content features in the image to be recognized for each line of text; splices the image features, text content features, and position features of each line of text to obtain multi-modal features of multiple lines of text; inputs the multi-modal features of multiple lines of text into a trained preset model, and outputs the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that through the technical solution provided by the embodiments of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content features in the image to be recognized for each line of text will be combined simultaneously to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition result can maintain semantic coherence.

[0121] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes computer instructions, and when the computer instructions run on an electronic device, enables the electronic device to execute the steps of the information recognition method described in the first aspect.

[0122] The technical solution provided by the embodiments of the present disclosure acquires an image to be recognized; respectively acquires the image features, text content features, and position features of the text content in the image to be recognized for each line of text; splices the image features, text content features, and position features of each line of text to obtain multi-modal features of multiple lines of text; inputs the multi-modal features of multiple lines of text into a trained preset model, and outputs the pointer positions corresponding to each line of text, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text. It can be seen that through the technical solution provided by the embodiments of the present disclosure, when performing text recognition on the image to be recognized, the image features, text content features, and position features of the text content in the image to be recognized for each line of text will be combined simultaneously to adjust the output positions of multiple lines of text, so that the output positions of multiple lines of text are more accurate, and thus the obtained text recognition results can maintain semantic coherence.

[0123] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, DSL (Digital Subscriber Line)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD (Digital Versatile Disc)), or a semiconductor medium (for example, an SSD (Solid State Disk)), etc.

[0124] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0125] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0126] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An information recognition method, characterized in that, it includes: Obtain an image to be recognized, where the image to be recognized includes multiple lines of text; Respectively obtain the image features, text content features of each line of text, and the position features of the text content features in the image to be recognized; Concatenate the image features, text content features, and position features of each line of text to obtain the multi-modal features of the multiple lines of text; Input the multi-modal features of the multiple lines of text into a trained preset model, and output the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text; Based on the pointer positions corresponding to each line of text in the image to be recognized, sort the multiple lines of text in the image to be recognized to obtain a sorting result; Display the multiple lines of text included in the image to be recognized according to the sorting result.

2. The method according to claim 1, characterized in that, The training process of the preset model includes: Obtain a training sample image, where the training sample image contains multiple lines of text; Extract the sample features of the training sample image, where the sample features include the image features, text content features of the multiple lines of text in the training sample image, and the position features of the text content features in the corresponding sample image; Input the sample features into the preset model to obtain the sample pointer positions corresponding to each line of text in the training sample image, calculate the loss value between the sample pointer positions and the labeled pointer positions corresponding to each line of text in the pre-labeled training sample image through the target loss function, and obtain the trained preset model when the loss value is less than the threshold.

3. The method according to any one of claims 1 to 2, characterized in that, The obtaining of the image features of each line of text includes: Extract the image features of each line of text through a convolutional neural network, where the image features include one or several combinations of text size features, text color features, texture features, and background features.

4. The method according to any one of claims 1 to 2, characterized in that, The image features, text content features, and the position features are respectively 256-dimensional feature vectors, and the multi-modal features are 256*3-dimensional feature vectors obtained by concatenating the image features, text content features, and the position features.

5. An information recognition device, characterized in that, it includes: An image acquisition module configured to execute the acquisition of an image to be recognized, where the image to be recognized includes multiple lines of text; A feature acquisition module configured to execute the acquisition of the image features, text content features of each line of text, and the position features of the text content features in the image to be recognized respectively; A feature concatenation module configured to execute the concatenation of the image features, text content features, and position features of each line of text to obtain the multi-modal features of the multiple lines of text; A position acquisition module configured to execute the input of the multi-modal features of the multiple lines of text into a trained preset model and output the pointer positions corresponding to each line of text respectively, where the pointer position corresponding to each line of text is used to represent the output position corresponding to that line of text; A text sorting module, configured to perform sorting on multiple lines of text in the image to be recognized based on the pointer positions corresponding to each line of text in the image to be recognized, so as to obtain a sorting result; A text display module, configured to perform displaying the multiple lines of text included in the image to be recognized according to the sorting result.

6. The apparatus according to claim 5, wherein, it further includes a feature training module, and the feature training module is specifically configured to perform: obtaining a training sample image, where the training sample image includes multiple lines of text; extracting sample features of the training sample image, where the sample features include image features, text content features of multiple lines of text in the training sample image, and position features of the text content features in the corresponding sample image; inputting the sample features into a preset model to obtain sample pointer positions corresponding to each line of text in the training sample image, calculating a loss value between the sample pointer positions and the labeled pointer positions corresponding to each line of text in the pre-labeled training sample image through a target loss function, and obtaining a trained preset model when the loss value is less than a threshold.

7. The apparatus according to any one of claims 5 to 6, wherein, the feature acquisition module is specifically configured to perform: extracting image features of each line of text through a convolutional neural network, where the image features include one or a combination of text size features, text color features, texture features, and background features.

8. The apparatus according to any one of claims 5 to 6, wherein, the image features, text content features, and the position features are respectively 256-dimensional feature vectors, and the multi-modal feature is a 256*3-dimensional feature vector obtained by splicing the image features, text content features, and the position features.

9. An electronic device, wherein, it includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the information recognition method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium, wherein, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the steps of the information recognition method according to any one of claims 1 to 4.

11. A computer program product, wherein, the computer program product includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is enabled to execute the steps of the information recognition method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Text structured extraction method, device and equipment and storage medium

    CN112001368A

  • OCR file format conversion method and system based on model optimization

    CN113065537A