An image recognition method, device, equipment, medium and product
By preprocessing the original image and extracting global contextual information using the Transformer model, the problems of insufficient contextual information capture and failure of long-distance dependency processing in existing image text recognition technologies are solved, achieving high-precision and robust text recognition and providing an intuitive display interface.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-16
AI Technical Summary
In existing technologies, image text recognition suffers from insufficient contextual information capture and failure in long-distance dependency processing when faced with images with complex backgrounds, long text sequences, and diverse styles. This results in high recognition errors, poor robustness, and weak scene adaptability.
By performing targeted preprocessing on the original image, extracting the visual features of the text, and combining the Transformer model to capture global contextual information, the image features and contextual information are fused together to improve recognition accuracy and robustness.
It improves the accuracy and robustness of text recognition in complex scenarios, reduces recognition errors, enhances the model's scene adaptability, and provides an intuitive visualization interface.
Smart Images

Figure CN122223728A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of image recognition technology, and more particularly to an image recognition method, apparatus, device, medium, and product. Background Technology
[0002] In the field of image text recognition, the extraction of text from images is a core component of applications such as document digitization, intelligent office work, and visual information retrieval.
[0003] In related technologies, image text recognition typically employs convolutional neural networks for feature extraction and recognition, extracting text content by performing global convolution operations on the entire image without discrimination. However, the accuracy of this method decreases with the complexity of the image scene. When faced with images with complex backgrounds, long text sequences, and diverse styles, it suffers from insufficient capture of contextual information and failure in handling long-distance dependencies, leading to high text recognition errors, poor robustness, and weak scene adaptability. Summary of the Invention
[0004] This disclosure addresses some deficiencies mentioned in the background art by providing an image recognition method, apparatus, device, medium, and product.
[0005] In a first aspect, embodiments of this disclosure provide an image recognition method, comprising: Receive the original image and perform image preprocessing on the original image to obtain the target image; Extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The visual features of the text and the global context information are fused to obtain the text recognition result of the original image; The text recognition results are displayed on the user's screen.
[0006] In one embodiment of the first aspect, the image preprocessing of the original image to obtain the target image includes: The original image is subjected to size unification processing to obtain an original image of a preset size; The original image of the preset size is converted to grayscale to obtain the grayscale original image; The original image after grayscale conversion is binarized to obtain the original image after binarization. The binarized original image is subjected to angle correction processing to obtain the corrected original image; The corrected original image is then subjected to denoising processing to obtain the target image.
[0007] In one embodiment of the first aspect, extracting image features from the target image to obtain text visual features includes: Image features are extracted from the target image using a convolutional neural network to obtain initial image features; The initial image features are serialized to obtain the text visual features.
[0008] In one embodiment of the first aspect, extracting global contextual information of the text content in the original image based on the text visual features includes: The visual features of the text are processed by the Transformer model to obtain global contextual information of the text content in the original image.
[0009] In one embodiment of the first aspect, processing the visual features of the text using the Transformer model to obtain global contextual information of the text content in the original image includes: The visual features of the text are processed by a Transformer encoder to obtain the initial contextual information of the text content in the original image. The initial context information is predicted by the Transformer decoder to obtain the global context information of the text content in the original image.
[0010] In one embodiment of the first aspect, after controlling the display of the text recognition result on the user's display interface, the method further includes: Receive editing instructions sent by the user; Based on the editing instructions, the text recognition results displayed on the display interface are modified to obtain the modified text recognition results; wherein, the editing instructions include, but are not limited to, modification instructions, text deletion instructions, and text addition instructions; The system receives a text export command sent by the user and transmits the modified text recognition result to the user's terminal.
[0011] In a second aspect, embodiments of this disclosure provide an image recognition device, comprising: The preprocessing module is used to receive the original image and perform image preprocessing on the original image to obtain the target image; The extraction module is used to extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The fusion module is used to fuse the visual features of the text and the global context information to obtain the text recognition result of the original image; The display module is used to control the display of the text recognition results on the user's display interface.
[0012] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of an image recognition method.
[0013] In a fourth aspect, a computer-readable storage medium is provided having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of an image recognition method.
[0014] In a fifth aspect, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of an image recognition method.
[0015] As will be described in detail below, an image recognition method, apparatus, device, medium, and product according to embodiments of the present disclosure are disclosed. By performing targeted image preprocessing on the original image to obtain the target image, interference information in the original image is effectively removed and the image shape is regularized, laying a high-quality data foundation for the accurate extraction of subsequent image features. Furthermore, by extracting the visual features of the text and the global contextual information of the text content in the original image in a layered manner, the limitations of related technologies that rely solely on convolutional neural networks to extract features indiscriminately and cannot capture global context are overcome. This approach preserves the morphological features at the visual level of the text and effectively mines long-distance dependencies of the text content, solving the problems of insufficient contextual information capture and failure of long-distance dependency processing in the background technology.
[0016] Building upon this foundation, this disclosure achieves text recognition results that combine visual morphological accuracy with contextual semantic plausibility through the fusion of image features and global contextual information. This improves text recognition accuracy in complex scenes, reduces recognition errors, and enhances the model's robustness and scene adaptability. Furthermore, the recognition results are visualized on the display interface, making them more intuitive and readily accessible. Attached Figure Description
[0017] Figure 1 A flowchart illustrating an image recognition method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating an image recognition method provided in this embodiment of the disclosure; Figure 3 A schematic diagram of an image recognition system provided in an embodiment of this disclosure; Figure 4 A schematic diagram of an image recognition device provided in an embodiment of this disclosure; Figure 5This is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0018] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present disclosure and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the drawings, not the entire structure.
[0019] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0020] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0021] Research has shown that in the field of image text recognition, the extraction of text from images is a core component of applications such as document digitization, intelligent office work, and visual information retrieval.
[0022] In related technologies, image text recognition typically employs convolutional neural networks for feature extraction and recognition, extracting text content by performing global convolution operations on the entire image without discrimination. However, the accuracy of this method decreases with the complexity of the image scene. When faced with images with complex backgrounds, long text sequences, and diverse styles, it suffers from insufficient capture of contextual information and failure in handling long-distance dependencies, leading to high text recognition errors, poor robustness, and weak scene adaptability.
[0023] Based on the above research, this disclosure provides an image recognition method. By performing targeted image preprocessing on the original image to obtain the target image, it effectively removes interference information and regularizes the image shape, laying a high-quality data foundation for the accurate extraction of subsequent image features. Furthermore, by extracting the visual features of the text and the global contextual information of the text content in the original image in a layered manner, it overcomes the limitations of related technologies that rely solely on convolutional neural networks to extract features indiscriminately and fail to capture global context. This method preserves the morphological features at the visual level of the text while effectively mining long-distance dependencies within the text content, solving the problems of insufficient contextual information capture and failure of long-distance dependency processing in the background technologies.
[0024] Building upon this foundation, this disclosure achieves text recognition results that combine visual morphological accuracy with contextual semantic plausibility through the fusion of image features and global contextual information. This improves text recognition accuracy in complex scenes, reduces recognition errors, and enhances the model's robustness and scene adaptability. Furthermore, the recognition results are visualized on the display interface, making them more intuitive and readily accessible.
[0025] To facilitate understanding of this embodiment, a detailed description of an image recognition method disclosed in this disclosure will be provided first. The execution subject of the image recognition method provided in this disclosure is generally an electronic device with a certain computing power. In some possible implementations, the image recognition method can be implemented by a processor calling computer-readable instructions stored in memory.
[0026] Figure 1 The diagram shows a flowchart of an image recognition method provided in an embodiment of this disclosure. The method includes steps S101 to S104, wherein: S101. Receive the original image and perform image preprocessing on the original image to obtain the target image.
[0027] In embodiments of this disclosure, the input image can be determined as the original image. Then, the original image can be optimized (i.e., image preprocessing) to obtain the target image.
[0028] Here, image preprocessing can improve the image quality of the original image and reduce interference in the original image, thus providing a higher quality target image for subsequent processing.
[0029] S102. Extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image.
[0030] In the embodiments of this disclosure, the target image can be processed by a pre-trained convolutional neural network (CNN) to obtain text visual features.
[0031] After determining the visual features of the text, the Transformer model can be used to extract the text and obtain the global contextual information of the text content in the original image.
[0032] S103. The visual features of the text and the global context information are fused to obtain the text recognition results of the original image.
[0033] In embodiments of this disclosure, the fusion strategy for the original image can be determined based on global context information.
[0034] For example, weight adjustment rules can be determined based on global context information, and these rules can be defined as the aforementioned fusion strategy. Specifically, the weight adjustment rules are used to adjust the proportion of visual feature weights corresponding to visual features and semantic feature weights corresponding to semantic features in the initial image features.
[0035] Among them, the visual adjustment weight + semantic adjustment weight = 1.
[0036] Then, the initial image features and global context information can be fused based on the fusion strategy to obtain the fused result of the original image.
[0037] Here, the fusion result can be processed through a context-aware mechanism to obtain the perception result; then, the perception result can be processed through an enhanced attention mechanism to obtain the text recognition result.
[0038] S104. Control the display of text recognition results on the user's display interface.
[0039] In the embodiments of this disclosure, the text corresponding to the text recognition result can be displayed on the display interface in a specific manner. This specific manner can be a display method that the user can distinguish.
[0040] For example, the text recognition results are highlighted in the original image, and the annotated original image is displayed on the user's screen.
[0041] In the embodiments of this disclosure, firstly, an original image is received and preprocessed to obtain a target image; secondly, image features of the target image are extracted to obtain text visual features, and global context information of the text content in the original image is extracted based on the text visual features; thirdly, the initial image features and global context information are fused to obtain the text recognition result of the original image; finally, the text recognition result is displayed on the user's display interface.
[0042] In the above embodiments, the target image is obtained by performing targeted image preprocessing on the original image, effectively removing interference information and regularizing the image shape, thus laying a high-quality data foundation for the accurate extraction of subsequent image features. Furthermore, by extracting the visual features of the text and the global contextual information of the text content in the original image in a layered manner, it overcomes the limitations of related technologies that rely solely on convolutional neural networks to extract features indiscriminately and fail to capture global context. This approach preserves the morphological features at the visual level of the text while effectively mining long-distance dependencies within the text content, solving the problems of insufficient contextual information capture and failure of long-distance dependency processing in the background technologies.
[0043] Building upon this foundation, this disclosure achieves text recognition results that combine visual morphological accuracy with contextual semantic plausibility through the fusion of image features and global contextual information. This improves text recognition accuracy in complex scenes, reduces recognition errors, and enhances the model's robustness and scene adaptability. Furthermore, the recognition results are visualized on the display interface, making them more intuitive and readily accessible.
[0044] In one optional implementation, image preprocessing is performed on the original image to obtain the target image, specifically including the following steps: First, the original image is resized to obtain an original image of the preset size; Secondly, the original image of the preset size is converted to grayscale to obtain the grayscale original image; Secondly, the original image after grayscale conversion is binarized to obtain the original image after binarization. Secondly, the original image after binarization is subjected to angle correction processing to obtain the corrected original image; Finally, the corrected original image is denoised to obtain the target image.
[0045] In embodiments of this disclosure, the original image can be normalized to adjust its size, resulting in an original image of a preset size. For example, the size of the original image can be uniformly adjusted to 1024×1024 pixels.
[0046] Here, after determining the original image of the preset size, the original image of the preset size can be converted into a grayscale original image using the OpenCV library to reduce unnecessary color information interference.
[0047] Here, after determining the original image after grayscale conversion, the original grayscale image can be binarized using the Otsu thresholding method to obtain the binarized original image, thereby enhancing the contrast between the text and the background in the original image.
[0048] Here, after determining the original image after binarization, the tilt angle of the original image after binarization can be detected by Hough transform, and the original image after binarization can be corrected based on the tilt angle to obtain the corrected original image.
[0049] Here, after determining the corrected original image, morphological operations can be used to denoise the corrected original image to obtain the target image.
[0050] In an optional implementation, image features of the target image are extracted to obtain text visual features, specifically including the following steps: First, image features are extracted from the target image using a convolutional neural network to obtain initial image features; Then, the initial image features are serialized to obtain the visual features of the text.
[0051] In the embodiments of this disclosure, image features can be extracted using a pre-trained convolutional neural network to obtain the initial visual features of the original image.
[0052] Next, global pooling can be performed on the initial visual features to convert them into a one-dimensional sequence. After determining the one-dimensional sequence, positional encoding can be added to the one-dimensional sequence to obtain the text visual features.
[0053] In one optional implementation, global contextual information of the text content in the original image is extracted based on the visual features of the text, specifically including the following steps: By processing the visual features of the text using the Transformer model, global contextual information of the text content in the original image can be obtained.
[0054] Here, the Transformer model is used to process the visual features of the text to obtain the global contextual information of the text content in the original image. The specific steps include the following: The Transformer encoder is used to process the visual features of the text to obtain the initial contextual information of the text content in the original image. The Transformer decoder predicts the initial context information to obtain the global context information of the text content in the original image.
[0055] In embodiments of this disclosure, the Transformer encoder includes a multi-head self-attention mechanism and a feedforward network, which can be used to extract initial contextual information of text content in the original image based on the visual features of the text.
[0056] Here, the Transformer decoder includes a self-attention mechanism and a feedforward neural network, which can be used to predict the initial context information to obtain the global context information of the text content in the original image.
[0057] Here, after determining the global context information, the global context information can be mapped to a character-level vocabulary.
[0058] In an optional implementation, after controlling the display of the text recognition results on the user's display interface, the following steps are included: First, receive editing instructions sent by the user; Secondly, based on editing commands, the text recognition results displayed on the screen are modified to obtain the modified text recognition results; among them, editing commands include, but are not limited to, modification commands, text deletion commands, and text addition commands; Finally, the system receives the text export command sent by the user and transmits the modified text recognition results to the user's terminal.
[0059] In the embodiments of this disclosure, the display interface is a visual operation interface that supports human-computer interaction. While presenting the text recognition results, the interface is configured with corresponding editing interaction entry points, including but not limited to text box editing areas, quick modification buttons, delete icons, and new input boxes.
[0060] Here, users can trigger editing commands through the above-mentioned interactive entry point based on the verification of the displayed results. The editing commands are collected by the user terminal and transmitted to the server that executes the image recognition method. The editing commands carry clear operation type identifiers and target operation location information.
[0061] The editing commands include, but are not limited to, modification commands, text deletion commands, and text addition commands: Modification commands are used to instruct the replacement of erroneous, ambiguous, or inaccurate text content in the displayed results, carrying the position coordinates of the text to be modified and the target replacement text; Text deletion commands are used to instruct the deletion of redundant, misidentified, or irrelevant text content in the displayed results, carrying the character range or position identifier of the text to be deleted; Text addition commands are used to instruct the addition of missing text content at a specified position in the displayed results, carrying the position index to be inserted and the newly added text content.
[0062] Here, the display interface can be synchronously refreshed in real time, and the changed text recognition results can be synchronously presented in the text display area of the display interface, and the highlighted annotations on the image can be updated to achieve real-time linkage, making it easy for users to check the changes immediately.
[0063] Here, the display interface is configured with a text export function entry. After the user has completed all the changes to the text recognition results and confirmed that they are correct, they can trigger the text export command through this entry.
[0064] Here, the display interface is equipped with a text deletion function. Users can delete the text recognition result after determining that it is no longer needed.
[0065] Reference Figure 2 The diagram shown is an overall flowchart of an image recognition method provided in this embodiment of the present disclosure, wherein: S10: Receive the original image input by the user.
[0066] S20. Perform image preprocessing on the original image to obtain the target image.
[0067] S30. Extract image features from the target image using a convolutional neural network to obtain the visual features of the text.
[0068] S40. Based on the visual features of text, extract the global contextual information of the text content in the original image using the Transformer model.
[0069] S50. Based on the target image, the visual features of the text and global context information are fused to obtain the text recognition result of the original image.
[0070] S60, Control the display of text recognition results on the user's display interface.
[0071] S70: Receive editing instructions sent by the user.
[0072] S80. Based on the editing command, modify the text recognition result displayed on the display interface to obtain the modified text recognition result.
[0073] S90: Receive the text export command sent by the user and transmit the modified text recognition result to the user's terminal.
[0074] In actual implementation, this embodiment has the following technical effects: (1) By proposing a Transformer-based text recognition algorithm framework, convolutional neural networks and Transformer models can be deeply integrated, and positional information of sequences can be preserved through positional encoding, thereby better handling long sequence data and solving the problems of long-term dependencies and insufficient contextual information. This algorithm framework not only retains the advantages of convolutional neural networks (CNNs) in capturing image detail features, but also effectively utilizes the high efficiency of Transformer models in handling long-distance dependencies, thus significantly improving the accuracy and stability of text recognition.
[0075] (2) The context-aware mechanism enables the model to dynamically adjust the fusion strategy according to the context of the current task, improving the model's adaptability and accuracy in complex scenarios. Then, the enhanced attention mechanism effectively combines low-level visual features and high-level semantic features. By fusing information at different levels, the recognition accuracy and robustness of the system are significantly improved, especially when dealing with complex scenarios and diverse text styles. This algorithm framework can not only focus on key information in different modalities, but also dynamically adjust the focus according to the context of the task.
[0076] (3) It provides a user-friendly and feature-rich interface that supports real-time recognition, result display and editing functions, improves the user experience, enables users to interact with the system conveniently, correct recognition errors in a timely manner, and enhances the practicality of the system.
[0077] Based on the same inventive concept, this disclosure also provides an image recognition system corresponding to the image recognition method. Since the principle of the system in this disclosure for solving the problem is similar to the image recognition method described above in this disclosure, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0078] Reference Figure 3 The diagram shown is a schematic representation of an image recognition system provided in this embodiment of the present disclosure, comprising: an image preprocessing unit 31, a context recognition unit 32, a fusion unit 33, and a visualization unit 34, wherein: The image preprocessing unit 31 includes: a size normalization subunit 311, a grayscale processing subunit 312, a binarization subunit 313, a correction subunit 314, and a noise reduction subunit 315.
[0079] Here, the size standardization subunit 311 is used to perform size standardization processing on the original image to obtain an original image of a preset size.
[0080] The grayscale processing subunit 312 is used to perform grayscale processing on the original image of the preset size to obtain the grayscale original image.
[0081] Binarization subunit 313 is used to perform binarization processing on the grayscale original image to obtain the binarized original image.
[0082] The correction subunit 314 is used to perform angle correction processing on the binarized original image to obtain the corrected original image.
[0083] The denoising subunit 315 is used to denoise the corrected original image to obtain the target image.
[0084] The context recognition unit 32 includes: a convolutional neural network layer 321, a serialization layer 322, a position encoding layer 323, a Transformer encoder 324, and a Transformer decoder 325.
[0085] Here, convolutional neural network layer 321 is used to extract image features through a pre-trained convolutional neural network to obtain the initial visual features of the original image.
[0086] Serialization layer 322 is used to perform global pooling on the initial visual features, converting the initial visual features into a one-dimensional sequence.
[0087] Position encoding layer 323 is used to add position encoding to a one-dimensional sequence to obtain visual features of text.
[0088] Transformer encoder 324 is used to process the visual features of text to obtain the initial context information of the text content in the original image.
[0089] The Transformer decoder 325 is used to predict the initial context information to obtain the global context information of the text content in the original image.
[0090] Here, the fusion unit 33 includes: fusion subunit 331, context-aware mechanism 332, and enhanced attention mechanism 333.
[0091] The fusion subunit 331 is used to fuse the initial image features and global context information to obtain the fusion result of the original image.
[0092] The context-aware mechanism 332 is used to process the fusion results to obtain the perception results.
[0093] An enhanced attention mechanism 333 is used to process the perception results to obtain the text recognition results.
[0094] Here, the visualization unit 34 includes: a result display subunit 341, a user editing subunit 342, and a saving subunit 343.
[0095] The result display subunit 341 is used to control the display of the text recognition results on the user's display interface.
[0096] User editing subunit 342 is used to receive editing instructions sent by the user; and based on the editing instructions, modify the text recognition results displayed on the display interface to obtain modified text recognition results.
[0097] The storage subunit 343 is used to receive the text export command sent by the user and transmit the modified text recognition result to the user's user terminal.
[0098] This embodiment obtains the target image by performing targeted image preprocessing on the original image, effectively removing interference information and regularizing the image shape, thus laying a high-quality data foundation for the accurate extraction of subsequent image features. Furthermore, by extracting the visual features of the text and the global contextual information of the text content in the original image in a layered manner, it overcomes the limitations of related technologies that rely solely on convolutional neural networks to extract features indiscriminately and fail to capture global context. This approach preserves the morphological features at the visual level of the text while effectively mining long-distance dependencies within the text content, solving the problems of insufficient contextual information capture and failure in long-distance dependency processing in the background technologies.
[0099] Building upon this foundation, this disclosure achieves text recognition results that combine visual morphological accuracy with contextual semantic plausibility through the fusion of image features and global contextual information. This improves text recognition accuracy in complex scenes, reduces recognition errors, and enhances the model's robustness and scene adaptability. Furthermore, the recognition results are visualized on the display interface, making them more intuitive and readily accessible.
[0100] Based on the same inventive concept, this disclosure also provides an image recognition device corresponding to the image recognition method. Since the principle of the device in this disclosure for solving the problem is similar to the image recognition method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0101] Reference Figure 4 The diagram shown is a schematic representation of an image recognition device according to an embodiment of this disclosure, comprising: a preprocessing module 41, an extraction module 42, a fusion module 43, and a display module 44; wherein: The preprocessing module is used to receive the original image and perform image preprocessing on the original image to obtain the target image; The extraction module is used to extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The fusion module is used to fuse the visual features of the text and the global context information to obtain the text recognition result of the original image; The display module is used to control the display of the text recognition results on the user's display interface.
[0102] This embodiment obtains the target image by performing targeted image preprocessing on the original image, effectively removing interference information and regularizing the image shape, thus laying a high-quality data foundation for the accurate extraction of subsequent image features. Furthermore, by extracting the visual features of the text and the global contextual information of the text content in the original image in a layered manner, it overcomes the limitations of related technologies that rely solely on convolutional neural networks to extract features indiscriminately and fail to capture global context. This approach preserves the morphological features at the visual level of the text while effectively mining long-distance dependencies within the text content, solving the problems of insufficient contextual information capture and failure in long-distance dependency processing in the background technologies.
[0103] Building upon this foundation, this disclosure achieves text recognition results that combine visual morphological accuracy with contextual semantic plausibility through the fusion of image features and global contextual information. This improves text recognition accuracy in complex scenes, reduces recognition errors, and enhances the model's robustness and scene adaptability. Furthermore, the recognition results are visualized on the display interface, making them more intuitive and readily accessible.
[0104] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0105] Corresponding to Figure 1 In addition to the image recognition method in this disclosure, an electronic device 500 is also provided, such as... Figure 5 The diagram shown is a structural schematic of an electronic device 500 provided in an embodiment of this disclosure, including: The system includes a processor 51, a memory 52, and a bus 53. The memory 52 stores execution instructions and includes main memory 521 and external memory 522. The main memory 521, also called internal memory, temporarily stores the computational data in the processor 51, as well as data exchanged with external memory such as a hard disk. The processor 51 exchanges data with the external memory 522 through the main memory 521. When the electronic device 500 is running, the processor 51 communicates with the memory 52 through the bus 53, causing the processor 51 to execute the following instructions: Receive the original image and perform image preprocessing on the original image to obtain the target image; Extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The visual features of the text and the global context information are fused to obtain the text recognition result of the original image; The text recognition results are displayed on the user's screen.
[0106] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0107] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0108] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.
[0109] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0110] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0111] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0112] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. An image recognition method, characterized in that, include: Receive the original image and perform image preprocessing on the original image to obtain the target image; Extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The visual features of the text and the global context information are fused to obtain the text recognition result of the original image; The text recognition results are displayed on the user's screen.
2. The method as described in claim 1, characterized in that, The step of preprocessing the original image to obtain the target image includes: The original image is subjected to size unification processing to obtain an original image of a preset size; The original image of the preset size is converted to grayscale to obtain the grayscale original image; The original image after grayscale conversion is binarized to obtain the original image after binarization. The binarized original image is subjected to angle correction processing to obtain the corrected original image; The corrected original image is then subjected to denoising processing to obtain the target image.
3. The method as described in claim 1, characterized in that, The step of extracting image features from the target image to obtain text visual features includes: Image features are extracted from the target image using a convolutional neural network to obtain initial image features; The initial image features are serialized to obtain the text visual features.
4. The method as described in claim 1, characterized in that, The step of extracting global contextual information of the text content in the original image based on the visual features of the text includes: The visual features of the text are processed using the Transformer model to obtain global contextual information of the text content in the original image.
5. The method as described in claim 4, characterized in that, The step of processing the visual features of the text using the Transformer model to obtain global contextual information of the text content in the original image includes: The visual features of the text are processed by a Transformer encoder to obtain the initial contextual information of the text content in the original image. The initial context information is predicted by the Transformer decoder to obtain the global context information of the text content in the original image.
6. The method as described in claim 1, characterized in that, After controlling the display of the text recognition result on the user's display interface, the method further includes: Receive editing instructions sent by the user; Based on the editing instructions, the text recognition results displayed on the display interface are modified to obtain the modified text recognition results; wherein, the editing instructions include, but are not limited to, modification instructions, text deletion instructions, and text addition instructions; The system receives a text export command sent by the user and transmits the modified text recognition result to the user's terminal.
7. An image recognition device, characterized in that, include: The preprocessing module is used to receive the original image and perform image preprocessing on the original image to obtain the target image; The extraction module is used to extract image features from the target image to obtain text visual features, and based on the text visual features, extract global context information of the text content in the original image; The fusion module is used to fuse the visual features of the text and the global context information to obtain the text recognition result of the original image; The display module is used to control the display of the text recognition results on the user's display interface.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 6.