An image information output method and related device
By obtaining and matching the field names and field contents in the image in image processing and matching according to the relative position constraint rules, the problem that different layout images require different templates to extract structured information is solved, and stronger generalization ability and applicability are achieved.
Patent Information
- Application Number
- CN202110180572.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Images of different layouts require different templates to effectively extract structured information, resulting in poor generalization ability and are not suitable for multiple information entry scenarios.
By obtaining multiple field names and multiple field contents in the target image, matching the extracted object according to the relative position constraint rules, obtaining the field name and field content that match successfully, thereby outputting structured information.
Without designing a corresponding template for each image, the extraction of structured information can be achieved, and the generalization ability is improved, and it is suitable for a variety of information entry scenarios.
Smart Images

Figure CN113569082B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image information output method and related devices. Background Art
[0002] With the development of science and technology, image recognition technology can be used to extract information from images. Taking bills as an example, image recognition technology can be used to extract information such as the title, amount, and time of the bill, thereby replacing the manual typing method to enter various bill information into the system, reducing the consumption of manpower and time.
[0003] For images such as bills with certain text arrangement rules, the field names and field contents included in the information to be extracted in the image have a structured relationship. For example, in a bill, "Time: January 1, 2020" is a structured information, where "Time" is the field name and "January 1, 2020" is its corresponding field content.
[0004] In the related technology, a corresponding information extraction template is formulated according to the text arrangement rules of the image. The template includes the position of the field content and the field name corresponding to the position. The field content in the image is extracted according to the position of the field content in the template, and it is combined with the corresponding field name to form a structured information, thereby completing the extraction of structured information in the image.
[0005] However, there are many formats of images. For example, taking bills alone as an example, they include value-added tax invoices, train tickets, taxi invoices, etc. Different formats have different text layout rules. Therefore, different templates are required for images of different formats to complete the extraction of structured information. The generalization ability is poor and it is not suitable for a variety of information entry scenarios. Summary of the invention
[0006] In order to solve the above technical problems, the present application provides an image information output method and related devices, which are used to solve the problem that different templates need to be formulated for images of different formats to complete the extraction of structured information.
[0007] The embodiments of the present application disclose the following technical solutions:
[0008] In one aspect, the present application provides a method for outputting image information, the method comprising:
[0009] Acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship;
[0010] Identifying a plurality of objects to be extracted in the target image, wherein the plurality of objects to be extracted include a plurality of field names and a plurality of field contents in the target image;
[0011] According to the relative position constraint rule, the multiple objects to be extracted are matched to obtain the field names and field contents that are successfully matched, wherein the relative position constraint rule is used to identify the relative positions of the field names and field contents having the structured relationship;
[0012] Output the field name and field content of the successful match.
[0013] On the other hand, the present application provides an image information output device, the device comprising: an acquisition unit, a recognition unit, a matching unit and an output unit;
[0014] The acquisition unit is used to acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship;
[0015] The identification unit is used to identify a plurality of objects to be extracted in the target image, wherein the plurality of objects to be extracted include a plurality of field names and a plurality of field contents in the target image;
[0016] The matching unit is used to match the multiple objects to be extracted according to the relative position constraint rule to obtain the field names and field contents that are successfully matched, and the relative position constraint rule is used to identify the relative positions of the field names and field contents that have the structured relationship;
[0017] The output unit is used to output the successfully matched field name and field content.
[0018] In another aspect, the present application provides a computer device, the device comprising a processor and a memory:
[0019] The memory is used to store program code and transmit the program code to the processor;
[0020] The processor is used to execute the method described in the above aspects according to the instructions in the program code.
[0021] On the other hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method described in the above aspects.
[0022] On the other hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above aspects.
[0023] It can be seen from the above technical solution that for a target image including multiple field names and field contents with a structured relationship, the field contents in the target image are no longer obtained according to the corresponding template, but multiple field names and multiple field contents in the target image are obtained, and the multiple field names and multiple field contents are used as objects to be extracted. According to the relative position constraint rule, the multiple objects to be extracted are matched to obtain the successfully matched field names and field contents, wherein the relative position constraint rule is used to identify the relative positions of the field names and field contents with a structured relationship, and the successfully matched field names and field contents can be obtained from the multiple objects to be extracted according to the relative positions between the multiple objects to be extracted, so as to output the successfully matched field names and field contents, and realize the output of the field names and field contents included in the target image according to the structured relationship. Therefore, by obtaining multiple field names and multiple field contents in the target image, and then matching the multiple field names and multiple field contents respectively, and obtaining the successfully matched field names and field contents according to the structured relationship, it is possible to realize the extraction of structured information without designing a corresponding template for each target image, thereby improving the generalization ability and being suitable for a variety of information entry scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 A schematic diagram of an application scenario of the image information output method provided in an embodiment of the present application;
[0026] Figure 2 A flowchart of an image information output method provided in an embodiment of the present application;
[0027] Figure 3 A schematic diagram of a target image provided for this application;
[0028] Figure 4 A schematic diagram of a target image recognition result provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of a multi-dimensional original feature fusion provided in an embodiment of the present application;
[0030] Figure 6 A schematic diagram of a feature interaction provided in an embodiment of the present application;
[0031] Figure 7A schematic diagram of a successfully matched field name and field content provided in an embodiment of the present application;
[0032] Figure 8 A schematic diagram of target image information output provided by an embodiment of the present application;
[0033] Fig. 9 A schematic diagram of a scenario of an image information output method provided in an embodiment of the present application;
[0034] Fig.10 A schematic diagram of an image information output device provided in an embodiment of the present application;
[0035] Fig.11 A schematic diagram of the structure of a server provided in an embodiment of the present application;
[0036] Fig.12 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The embodiments of the present application are described below in conjunction with the accompanying drawings.
[0038] In view of the fact that the related art extracts structured information from images according to templates, the generalization ability is poor and it is not suitable for various information entry scenarios. This application proposes a method and a related device for outputting image information, which can output structured information in images without templates and is suitable for various information entry scenarios.
[0039] The image information output method provided in the embodiment of the present application is based on artificial intelligence. Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0040] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0041] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned natural language processing, machine learning / deep learning and other directions. For example, it may involve text preprocessing (Text preprocessing) and semantic understanding (Semantic understanding) in natural language processing (NLP), and it may also involve deep learning (Deep Learning) in machine learning (Machine learning, ML), including various types of artificial neural networks (Artificial Neural Network, ANN).
[0042] The image information output method provided in the present application can be applied to image information output devices with data processing capabilities, such as terminal devices and servers. The terminal device can specifically be a smart phone, a desktop computer, a laptop computer, a tablet computer, a smart speaker, a smart watch, etc., but is not limited thereto; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.
[0043] The image information output device may have the ability to implement natural language processing, which is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language used by people in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies. In an embodiment of the present application, the text processing device can process the text through technologies such as text preprocessing and semantic understanding in natural language processing.
[0044] The image information output device may have machine learning capabilities. Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks.
[0045] The artificial intelligence model used in the image information output method provided in the embodiment of the present application mainly involves the application of neural networks, which are used to identify multiple objects to be extracted in the target image and match multiple objects to be extracted.
[0046] In order to facilitate understanding of the technical solution of the present application, the entity type determination method provided in the embodiment of the present application is introduced below in combination with actual application scenarios.
[0047] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of the image information output method provided in an embodiment of the present application. Figure 1 In the application scenario shown, the aforementioned image information output device is a server 100 .
[0048] The server 100 obtains a target image, which is an image with certain text arrangement rules, including multiple field names and field contents with structural relationships, and the field names and the field contents corresponding to the field names can constitute a structured information. Figure 1 In the application scenario shown, the target image is the ID card image of Zhang San, and "Name Zhang San" is a structured information, where "Name" is the field name and "Zhang San" is the field content.
[0049] The server 100 identifies multiple objects to be extracted in the target image, where the objects to be extracted are structured information in the target image, including multiple field names and multiple field contents in the target image. Figure 1 In the application scenario shown, the server 100 locates all field names in the target image through a solid rectangular detection box and extracts corresponding position features, and locates all field contents in the target image through a dotted rectangular detection box and extracts corresponding position features.
[0050] The server 100 matches the identified multiple field names and multiple field contents according to the relative position constraint rule to obtain the successfully matched field names and field contents. The relative position constraint rule is used to identify the relative positions of the field names and field contents with a structured relationship, and can obtain the successfully matched field names and field contents from the multiple objects to be extracted according to the relative positions between the multiple objects to be extracted, thereby outputting the successfully matched field names and field contents, and realizing the output of the field names and field contents included in the image according to the structured relationship.
[0051] exist Figure 1In the application scenario shown, the field name "name" is matched with the corresponding field content as an example. Through the relative position constraint rule, "Zhang San" has a structural relationship with the field name "name" in multiple field contents, and the field name "name" and the field content "Zhang San" can be used as the field name and field content that are successfully matched. However, "male" does not have a structural relationship with the field name "name" in multiple field contents, and the field name "name" and the field content "male" fail to match.
[0052] After matching the identified multiple field names and multiple field contents, the server 100 outputs the successfully matched field names and field contents. For example, "Name: Zhang San" is a piece of structured information, namely, the successfully matched field name and field content.
[0053] Therefore, by obtaining multiple field names and multiple field contents in the target image, and then matching the multiple field names and multiple field contents in the target image respectively, the successfully matched field names and field contents are obtained according to the structured relationship, thereby realizing the extraction of structured information without designing a corresponding template for each image, thereby improving the generalization ability and being suitable for a variety of information entry scenarios.
[0054] In conjunction with the accompanying drawings, an image information output method provided by an embodiment of the present application is introduced below with a server as an image information output device.
[0055] See also Figure 2 , Figure 2 This is a flow chart of an image information output method provided in an embodiment of the present application. Figure 2 As shown, the image information output method includes the following steps:
[0056] S201: Acquire a target image.
[0057] The target image is an image with certain text arrangement rules, including multiple field names and field contents with structured relationships. One piece of structured information includes a field name and field content corresponding to the field name.
[0058] The embodiment of the present application does not specifically limit the type of target image. For example, the target image may include one or more of a license image, a certificate image, a document image, and a bill image. The technical solution of the present application can be applied to, for example, the structured information extraction of the aforementioned multiple images, without the need to design a corresponding template for each image, and has a strong generalization capability. After the structured information in multiple images is extracted, the information can be automatically entered into the system.
[0059] For the convenience of explanation, the target image is taken as a bill image in the embodiment of the present application. Figure 3 , which is a schematic diagram of a target image provided by this application. Figure 3 As shown, this figure is an image of the billing receipt for Zhang San when he was hospitalized in the Central Medical Hospital.
[0060] S202: Identify multiple objects to be extracted in the target image.
[0061] The target image includes multiple field names and multiple field contents. In the related technology, the field content is only identified through the template corresponding to the target image. In the embodiment of the present application, not only the field content but also the field name is identified. The field names and field contents included in the target image are all objects to be extracted.
[0062] The embodiment of the present application does not specifically limit the method of identifying the object to be extracted, and is described below using a labeled text detection network as an example.
[0063] In the related art, a trained text detection network is used to identify the field content in an image. When training the text detection network, the position of the field content recorded in the template corresponding to the image is used as a label. For example, the template corresponding to the target image is the target template, and the target template includes m field content positions, correspondingly setting m labels, one label represents the position of one field content, thereby training the text detection network corresponding to the target template. In other words, the text detection network trained according to the target template is only applicable to the target image, and has poor generalization ability, which not only causes model redundancy but also increases the workload of the customized model.
[0064] However, when training a text detection network, the embodiment of the present application no longer sets multiple field content labels at different positions according to the corresponding template, but sets two types of labels, namely field name labels and field content labels, so that the trained text detection network can distinguish between field names and field contents in an image, and the trained text detection network model is no longer limited to one type of image, but can be used for multiple images, has universality, strong generalization ability, and reduces the workload caused by customization.
[0065] As a possible implementation method, when training the text detection network, the background area can also be set as the third type of label, that is, the labels in the labeled text detection network can include not only field name class labels and field content class labels, but also background area class labels. Among them, the background area class label is information that does not need to be extracted, for example, Figure 3 The information on the middle right side such as "First Copy" and the title "Central Medical Hospitalization Fee Receipt" are all information that does not need to be extracted, that is, the background area.
[0066] By combining the background area class labels, the text detection network can better learn the characteristics of field names and field contents, thereby avoiding the jitter of the text detection network caused by the interference of the background area and enhancing the robustness of the text detection network.
[0067] See also Figure 4 , which is a schematic diagram of a target image recognition result provided by an embodiment of the present application. Figure 4 In the example, a text detection network model is used to detect the field name, field content, and background area in the fee receipt, where the solid rectangular detection box is the field name, the dotted rectangular detection box is the field content, and the solid circular detection box is the background area.
[0068] After the field name, field content and background area in the target image are detected by the text detection network, the detection result can be used as input to obtain the text content corresponding to the field name, field content and background area through the recognition network.
[0069] As a possible implementation method, a network model can also be trained, through which the text content in the target image can be first identified, and then the type corresponding to the text content can be identified through the text content, such as field name, field content and background area.
[0070] By using the labeled text detection network provided in the embodiment of the present application to identify the field name and field content in the target image and to identify the content thereof, the problem of field adhesion can be effectively solved.
[0071] For example, in Figure 4 The "male" and "medical insurance type" are stuck together. If the method in the related technology is adopted, it is easy to extract "male" and "medical insurance type" together as the field content corresponding to the field name "gender:". However, with the technical solution of this application, since "male" will be recognized as field content and "medical insurance type:" will be recognized as field name, the two will not be stuck together, thus solving the problem of field sticking.
[0072] S203: Match multiple objects to be extracted according to the relative position constraint rule to obtain the field names and field contents of successful matches.
[0073] After obtaining multiple field names and field contents in the target image, the field names and field contents in the target image are matched according to relative position constraint rules, wherein the relative position constraint rules are used to identify the relative positions of field names and field contents having a structured relationship. According to the relative positions between the field names and field contents in the target image, successfully matched field names and field contents can be obtained from multiple objects to be extracted, thereby outputting successfully matched field names and field contents, and the successfully matched field names and field contents can constitute structured information.
[0074] The present application embodiment does not specifically limit the method of obtaining the relative position of the field name and field content. Figure 4The position features of the detection box in are used as an example to illustrate.
[0075] By extracting original features of each detection box (including at least the detection box corresponding to the field name and the detection box corresponding to the field content), such as position features, and comparing the position features of each detection box, the relative position of the field name and the field content is obtained.
[0076] As a possible implementation method, not only the position features of the object to be extracted can be extracted, but also the original features of other dimensions of the object to be extracted can be extracted, and the original features of multiple dimensions of the object to be extracted can be fused. Compared with the detection frame matching only through the position features, considering the matching relationship between the objects to be extracted from multiple aspects can effectively solve the problem of field offset and improve the accuracy of matching between detection frames.
[0077] Next, continue to combine Figure 4 , taking the original features as position features, image features and semantic features as examples, the process of extracting the features of the object to be extracted is explained.
[0078] To facilitate subsequent processing, the size of the target image can be normalized. For example, the resolution of the target image can be scaled to 512*512 to obtain a normalized image. A plane coordinate system is established to obtain the scaling factor scale_x of the target image in the x direction and the scaling factor scale_y of the target image in the y direction. The normalized target image is then used as input to obtain a feature map with a resolution of 128*128 and 64 channels through the cropped ResNet residual network. This feature map is used to characterize the texture, color, and shape of the target image.
[0079] (1) Image features.
[0080] Figure 4 The coordinates of each detection frame in is (xi, yi, wi, hi), where xi represents the position of the ith detection frame on the x-axis, yi represents the position of the ith detection frame on the y-axis, wi represents the width of the ith detection frame, and hi represents the height of the ith detection frame. After scaling the coordinates of each detection frame (xi, yi, wi, hi) by (128 / 512)*scale_x and (128 / 512)*scale_y, we get the coordinates of each detection frame on the bottom-level feature map (xi', yi', wi', hi'). Finally, we take the feature at the position (xi'+wi' / 2, yi'+hi' / 2) in the feature map as the feature of the ith detection frame, and get Figure 4 The image feature set of n detection boxes in , each image feature dimension is 64.
[0081] (2) Location characteristics.
[0082] Will Figure 4 The coordinates (xi, yi, wi, hi) of each detection box in are used as position features to obtain Figure 4 The position feature set of n detection boxes in , each position feature dimension is 4.
[0083] (3)Semantic features.
[0084] Use text embedding to generate word vectors for the text information in each detection box, and get Figure 4 The semantic feature set of n detection boxes in , and the dimension of each semantic feature is 64.
[0085] After extracting original features of multiple different dimensions of the object to be extracted, the original features of the multiple dimensions are fused as features of the object to be extracted. The following continues to illustrate by taking the extraction of image features, position features and semantic features of the detection frame as an example.
[0086] See also Figure 5 , which is a schematic diagram of a multi-dimensional original feature fusion provided by an embodiment of the present application. Figure 5 In the figure, 3 layers are taken as an example to represent the multi-layer feature map of the detection frame, which are represented as 5011, 5012 and 5013 respectively, where 5013 represents the feature map of the bottom layer of the detection frame, and the image feature 502 in the feature map of this layer is extracted, and the original features of the three dimensions of the position feature 503 and the semantic feature 504 are fused to be used as the feature 505 of the detection frame. By repeating the above process, the features of n detection frames in the target image can be obtained, and the features of each object to be extracted are 132 dimensions.
[0087] The relative position constraint rule is used to identify the relative position of field names and field contents with a structured relationship. It pays more attention to the field names and field contents that constitute a piece of structured information in the image, and pays less attention to the features between multiple pieces of structured information. When there is an offset between local texts in the target image or there is obvious ambiguity between fields, it is easy to fall into local optimality, reducing the accuracy of matching field names and field contents.
[0088] Based on this, in order to avoid the above problems, the features of the object to be extracted and the features of other objects to be extracted adjacent to the object to be extracted can be interacted to obtain interactive features of the object to be extracted. The following is an example of a feature interaction model.
[0089] See also Figure 6 , which is a schematic diagram of a feature interaction provided by an embodiment of the present application. Figure 6 As shown in the upper left corner, the features of the object to be extracted are input into the feature interaction model, and the feature interaction model can interact the features of the object to be extracted with the features of other objects to be extracted around it.
[0090] Taking the aforementioned rectangular detection box as an example, obtain the coordinates of each vertex of the i-th rectangular detection box. Taking a vertex as an example, find the k vertices closest to the vertex, and fuse the features of the k nearest vertices into the vertex, thereby strengthening the connection between the vertices, and then making the detection boxes connected, improving the global features. Then, the features of each vertex are fused together through, for example, a fully connected neural network (FCN) model, so that the features of each vertex are reduced from 132 dimensions to 64 dimensions, completing the feature interaction process. According to the relative position constraint rules and the interactive features of the objects to be extracted, multiple objects to be extracted are matched to obtain the field names and field contents of successful matches.
[0091] As a possible implementation method, in the matching process, any two rectangular detection frames can be matched to obtain n*n matching combinations (including the matching of the i-th rectangular detection frame and the i-th rectangular detection frame), and the n*n matching combinations are used as inputs to the classification model, and the classification model is used to predict whether any two rectangular detection frames match. During the prediction, it is necessary to predict whether each pair of combinations matches, because it is necessary to obtain the matching results of all rectangular detection frames. For example, first concatenate the features of any two rectangular detection frames, changing from n*64 dimensions to n*n*128 dimensions, such as Figure 6 As shown below, an adjacency matrix representing the n rectangular detection boxes in the target image is obtained, and a classification model is used to determine whether the combination represented by the adjacency matrix has a matching relationship.
[0092] As a possible implementation method, when training the classification model, Monte Carlo sampling can be used to randomly select pairs of rectangular detection boxes for training. Since only a part of the detection boxes instead of all the detection boxes are used for training, the efficiency and randomness of the training can be improved.
[0093] After obtaining whether the field name and field content match, the following describes the process of completing the matching of the field name and field content, see S2031-S2034.
[0094] S2031: Acquire at least one second object to be extracted that matches the first object to be extracted according to the relative position constraint rule.
[0095] Taking a rectangular detection box as an example, the rectangular detection box is the first object to be extracted. The first object to be extracted can be a solid rectangular detection box corresponding to the field name, or a dotted rectangular detection box corresponding to the field content. This application does not make specific limitations.
[0096] Obtain at least one second object to be extracted that matches the first object to be extracted. The second object to be extracted can be a solid rectangular detection box corresponding to the field name, or a dotted rectangular detection box corresponding to the field content. This application does not make specific limitations.
[0097] S20332: Put the first object to be extracted and the second object to be extracted into a set.
[0098] Put the rectangular detection frames that have matching relationships with each other into the same set. For example, traverse the adjacency matrix and select the rectangular detection frame that has not been transformed (not transformed from the adjacency matrix to field content or field name) as the first object to be extracted; continue to traverse the adjacency matrix, select the second object to be extracted that has not been transformed, and the second object to be extracted can be multiple, put the first object to be extracted and the second object to be extracted into a set, and change the status of the first object to be extracted and the second object to be extracted to transformed; repeat the above steps until the status of all matrix detection frames in the adjacency matrix is transformed.
[0099] S2033: If there are multiple field names in the set, determine the arrangement order of the field names in the set according to the relative positions of the field names in the set; if there are multiple field contents in the set, determine the arrangement order of the field contents in the set according to the relative positions of the field contents in the set.
[0100] by Figure 4 For example, according to the relative position constraint rule, the field name "Hospitalization time:" and the two field contents "March 14, 2016" and "March 21, 2016" will be put into the same set. In this set, there are two field contents. According to the relative position of the two field contents, it can be determined that the arrangement order of the field contents in the set should be that the field content "March 14, 2016" is before the field content "March 21, 2016".
[0101] S2034: Sort the first object to be extracted and the second object to be extracted according to the arrangement order to obtain the field name and field content that are successfully matched.
[0102] Continuing with the example in S2033, the order of the field name "Hospitalization Time:" and the two field contents "March 14, 2016" and "March 21, 2016" should be "Hospitalization Time:", "March 14, 2016", and "March 21, 2016".
[0103] See also Figure 7, this figure is a schematic diagram of a successfully matched field name and field content provided in an embodiment of the present application. In order to clearly illustrate the matching relationship between the field name and the field content, the successfully matched field name and field content are rendered with the same rectangular frame. For example, the field name "name" and the field content "Zhang San" are a piece of structured information, and the field name "gender" and the field content "male" are a piece of structured information. The two pieces of structured information are rendered with different rectangular frames.
[0104] Therefore, by putting multiple objects to be extracted that have a matching relationship with each other into a set, and then re-sorting them according to their positional relationship, and finally outputting them, the problem of low matching accuracy caused by field wrapping and many-to-many can be effectively solved.
[0105] S204: Output the successfully matched field names and field contents.
[0106] See also Figure 8 , which is a schematic diagram of target image information output provided by an embodiment of the present application. Figure 8 In the , the successfully matched field names and field contents are displayed on the right side of the target avatar for users to check. The successfully matched field names and field contents can be entered into the system.
[0107] It can be seen from the above technical solution that for a target image including multiple field names and field contents with a structured relationship, the field contents in the target image are no longer obtained according to the corresponding template, but multiple field names and multiple field contents in the target image are obtained, and the multiple field names and multiple field contents are used as objects to be extracted. According to the relative position constraint rule, the multiple objects to be extracted are matched to obtain the successfully matched field names and field contents, wherein the relative position constraint rule is used to identify the relative positions of the field names and field contents with a structured relationship, and the successfully matched field names and field contents can be obtained from the multiple objects to be extracted according to the relative positions between the multiple objects to be extracted, so as to output the successfully matched field names and field contents, and realize the output of the field names and field contents included in the target image according to the structured relationship. Therefore, by obtaining multiple field names and multiple field contents in the target image, and then matching the multiple field names and multiple field contents respectively, and obtaining the successfully matched field names and field contents according to the structured relationship, it is possible to realize the extraction of structured information without designing a corresponding template for each target image, thereby improving the generalization ability and being suitable for a variety of information entry scenarios.
[0108] In order to better understand the image information output method provided in the embodiment of the present application, Fig. 9 The image information output method provided in the embodiment of the present application is described.
[0109] See also Fig. 9 , which is a scene diagram of an image information output method provided by an embodiment of the present application. Fig. 9 In the illustrated scenario, the target image may be any image having multiple pieces of structured information.
[0110] S901: Text detection and recognition.
[0111] The labeled text detection network locates the field name, field content, and background area in the target image through the detection box, and identifies the information content in each detection box.
[0112] S902: Feature extraction.
[0113] The image features, position features and semantic features of each detection box are extracted through Graph Neural Networks (GNN).
[0114] S903: Feature fusion.
[0115] For each detection frame, the image features, position features and semantic features are fused to obtain the fused features of each detection frame.
[0116] S904: Feature interaction.
[0117] The nearest detection box of each detection box is found through the K-Nearest Neighbor algorithm (KNN), and they are integrated together through FCN to obtain the adjacency matrix.
[0118] S905: Matching relationship prediction.
[0119] The adjacency matrix is input into the classification model, and the classification model is used to predict whether every two detection boxes have a matching relationship.
[0120] S906: Adjacency matrix analysis.
[0121] The adjacency matrix is parsed using the S2031-S2034 method to obtain the successfully matched field names and field contents, and output them as structured information. This is suitable for various information entry scenarios, saving labor costs and improving enterprise efficiency.
[0122] The image information output method provided in the embodiment of the present application is evaluated by constructing a data set for measuring the accuracy of structured information extraction. As shown in Table 1, the recall rate is 92.49% and the accuracy rate is 93.77%.
[0123] Table 1 Structured information extraction indicators
[0124] Recall Accuracy Structured Information Extraction 92.49% 93.77%
[0125] Among the evaluation indicators, recall rate refers to the proportion of correct fields extracted from the object to be extracted in the information existing in the data set; accuracy rate refers to the proportion of correct information extracted from the object to be extracted in the predicted information.
[0126] In view of the image information output method provided in the above embodiment, the embodiment of the present application also provides an image information output device.
[0127] See also Fig.10 , which is a schematic diagram of an image information output device provided in an embodiment of the present application. Fig.10 As shown, the image output device 1000 includes: an acquisition unit 1001, a recognition unit 1002, a matching unit 1003 and an output unit 1004;
[0128] The acquisition unit 1001 is used to acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship;
[0129] The identification unit 1002 is used to identify multiple objects to be extracted in the target image, where the multiple objects to be extracted include multiple field names and multiple field contents in the target image;
[0130] The matching unit 1003 is used to match the multiple objects to be extracted according to the relative position constraint rule to obtain the field names and field contents that are successfully matched, and the relative position constraint rule is used to identify the relative positions of the field names and field contents that have the structured relationship;
[0131] The output unit 1004 is used to output the successfully matched field name and field content.
[0132] As a possible implementation manner, the matching unit 1003 is configured to:
[0133] According to the relative position constraint rule, obtaining at least one second object to be extracted that matches the first object to be extracted;
[0134] Putting the first object to be extracted and the second object to be extracted into a set;
[0135] If there are multiple field names in the set, the arrangement order of the field names in the set is determined according to the relative positions of the field names in the set; if there are multiple field contents in the set, the arrangement order of the field contents in the set is determined according to the relative positions of the field contents in the set;
[0136] The first to-be-extracted objects and the second to-be-extracted objects are sorted according to the arrangement order to obtain field names and field contents that have been matched successfully.
[0137] As a possible implementation manner, the matching unit 1003 is configured to:
[0138] Interacting the features of the object to be extracted and the features of other objects to be extracted that are adjacent to the position of the object to be extracted to obtain interactive features of the object to be extracted;
[0139] According to the relative position constraint rules and the interactive features of the objects to be extracted, the multiple objects to be extracted are matched to obtain the field names and field contents of successful matches.
[0140] As a possible implementation manner, before the matching unit 1003 interacts the features of the object to be extracted and the features of other objects to be extracted that are adjacent to the position of the object to be extracted to obtain the interactive features of the object to be extracted, the device 1000 is further used to:
[0141] Extracting original features of the object to be extracted, wherein the original features include at least one of image features, position features and semantic features;
[0142] If the original features of the object to be extracted include multiple types, the original features of the object to be extracted are fused to obtain the features of the object to be extracted.
[0143] As a possible implementation manner, the identification unit 1002 is configured to:
[0144] A plurality of objects to be extracted in the target image are identified according to a labeled text detection network, wherein the labels in the labeled text detection network include field name class labels and field content class labels.
[0145] As a possible implementation manner, the labels in the labeled text detection network also include background area class labels.
[0146] As a possible implementation manner, the target image is at least one of a license image, a certificate image, a document image, and a bill image.
[0147] The image information output device provided in the above embodiment, for a target image including multiple field names and field contents with a structured relationship, no longer obtains the field contents in the target image according to the corresponding template, but obtains multiple field names and multiple field contents in the target image, takes the multiple field names and multiple field contents as objects to be extracted, matches the multiple objects to be extracted according to the relative position constraint rule, and obtains the successfully matched field names and field contents, wherein the relative position constraint rule is used to identify the relative positions of the field names and field contents with a structured relationship, and can obtain the successfully matched field names and field contents from the multiple objects to be extracted according to the relative positions between the multiple objects to be extracted, thereby outputting the successfully matched field names and field contents, and outputting the field names and field contents included in the target image according to the structured relationship. Thus, by obtaining multiple field names and multiple field contents in the target image, and then matching the multiple field names and multiple field contents respectively, and obtaining the successfully matched field names and field contents according to the structured relationship, it is possible to extract structured information without designing a corresponding template for each target image, thereby improving the generalization ability and being applicable to a variety of information entry scenarios.
[0148] The embodiment of the present application also provides a computer device. The computer device provided by the embodiment of the present application will be introduced from the perspective of hardware instantiation.
[0149] See also Fig.11 , Fig.11 14 is a schematic diagram of a server structure provided in an embodiment of the present application. The server 1400 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1422 (for example, one or more processors) and memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 may be short-term storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1422 may be configured to communicate with the storage medium 1430 to execute a series of instruction operations in the storage medium 1430 on the server 1400.
[0150] The server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0151] The steps performed by the server in the above embodiment can be based on the Fig.11 The server structure shown.
[0152] The CPU 1422 is used to execute the following steps:
[0153] Acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship;
[0154] Identifying a plurality of objects to be extracted in the target image, wherein the plurality of objects to be extracted include a plurality of field names and a plurality of field contents in the target image;
[0155] According to the relative position constraint rule, the multiple objects to be extracted are matched to obtain the field names and field contents that are successfully matched, wherein the relative position constraint rule is used to identify the relative positions of the field names and field contents having the structured relationship;
[0156] Output the field name and field content of the successful match.
[0157] Optionally, the CPU 1422 may also execute the method steps of any specific implementation of the image information output method in the embodiments of the present application.
[0158] With respect to the image information output method described above, an embodiment of the present application further provides a terminal device for outputting image information, so that the above-mentioned image information output method can be implemented and applied in practice.
[0159] See also Fig.12 , Fig.12 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. For the sake of convenience, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), etc., taking the terminal device as a mobile phone as an example:
[0160] Fig.12 FIG. 1 is a block diagram showing a partial structure of a mobile phone related to a terminal device provided in an embodiment of the present application. Fig.12The mobile phone includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a wireless fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590. Those skilled in the art will understand that Fig.12 The mobile phone structure shown in the figure does not constitute a limitation on the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0161] Combine the following Fig.12 A detailed introduction to the various components of the mobile phone:
[0162] The RF circuit 1510 can be used for receiving and sending signals during the process of sending and receiving information or making calls. In particular, after receiving the downlink information of the base station, it is sent to the processor 1580 for processing; in addition, the designed uplink data is sent to the base station. Usually, the RF circuit 1510 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (Low Noise Amplifier, referred to as LNA), a duplexer, etc. In addition, the RF circuit 1510 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0163] The memory 1520 can be used to store software programs and modules. The processor 1580 implements various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1520. The memory 1520 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 1520 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0164] The input unit 1530 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 1530 may include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 1531 or near the touch panel 1531 using any suitable object or accessory such as a finger, stylus, etc.), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1531 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1580, and can receive and execute the command sent by the processor 1580. In addition, the touch panel 1531 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic waves. In addition to the touch panel 1531, the input unit 1530 may also include other input devices 1532. Specifically, other input devices 1532 may include but are not limited to one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, and the like.
[0165] The display unit 1540 may be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1540 may include a display panel 1541. Optionally, the display panel 1541 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1531 may cover the display panel 1541. When the touch panel 1531 detects a touch operation on or near it, it is transmitted to the processor 1580 to determine the type of touch event. Subsequently, the processor 1580 provides a corresponding visual output on the display panel 1541 according to the type of touch event. Although in Fig.12 In the embodiment, the touch panel 1531 and the display panel 1541 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 1531 and the display panel 1541 can be integrated to realize the input and output functions of the mobile phone.
[0166] The mobile phone may also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1541 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1541 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.
[0167] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the mobile phone. The audio circuit 1560 can transmit the received audio data to the speaker 1561 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and converted into audio data, and then the audio data is output to the processor 1580 for processing, and then sent to another mobile phone through the RF circuit 1510, or the audio data is output to the memory 1520 for further processing.
[0168] WiFi is a short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse web pages and access streaming media through the WiFi module 1570. It provides users with wireless broadband Internet access. Fig.12 A WiFi module 1570 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.
[0169] The processor 1580 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. It executes various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1520 and calling data stored in the memory 1520. Optionally, the processor 1580 may include one or more processing units; preferably, the processor 1580 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1580.
[0170] The mobile phone also includes a power supply 1590 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.
[0171] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.
[0172] In the embodiment of the present application, the memory 1520 included in the mobile phone can store program codes and transmit the program codes to the processor.
[0173] The processor 1580 included in the mobile phone can execute the image information output method provided in the above embodiment according to the instructions in the program code.
[0174] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program, wherein the computer program is used to execute the image information output method provided in the above embodiment.
[0175] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the image information output method provided in various optional implementations of the above aspects.
[0176] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the above-mentioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the above-mentioned storage medium can be at least one of the following media: read-only memory (English: read-only memory, abbreviated: ROM), RAM, magnetic disk or optical disk, etc. Various media that can store program codes.
[0177] It should be noted that each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, in which the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0178] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for outputting image information, It is characterized in that The method comprises: Acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship; Identifying a plurality of objects to be extracted in the target image, wherein the plurality of objects to be extracted include a plurality of field names and a plurality of field contents in the target image; According to the relative position constraint rule, the multiple objects to be extracted are matched to obtain the field names and field contents that are successfully matched, specifically including: interacting the features of the object to be extracted and the features of other objects to be extracted that are adjacent to the position of the object to be extracted to obtain the interactive features of the object to be extracted; according to the relative position constraint rule and the interactive features of the object to be extracted, the multiple objects to be extracted are matched to obtain the field names and field contents that are successfully matched; the relative position constraint rule is used to identify the relative positions of the field names and field contents that have the structured relationship; Output the field name and field content of the successful match.
2. The method according to claim 1, It is characterized in that The step of matching the plurality of objects to be extracted according to the relative position constraint rule to obtain the field names and field contents of successful matches also includes: According to the relative position constraint rule, obtaining at least one second object to be extracted that matches the first object to be extracted; Putting the first object to be extracted and the second object to be extracted into a set; If there are multiple field names in the set, the arrangement order of the field names in the set is determined according to the relative positions of the field names in the set; if there are multiple field contents in the set, the arrangement order of the field contents in the set is determined according to the relative positions of the field contents in the set; The first to-be-extracted objects and the second to-be-extracted objects are sorted according to the arrangement order to obtain field names and field contents that have been matched successfully.
3. The method according to claim 1, It is characterized in that Before the feature of the object to be extracted and the features of other objects to be extracted adjacent to the position of the object to be extracted are interacted to obtain the interactive features of the object to be extracted, the method further includes: Extracting original features of the object to be extracted, wherein the original features include at least one of image features, position features and semantic features; If the original features of the object to be extracted include multiple types, the original features of the object to be extracted are fused to obtain the features of the object to be extracted.
4. The method according to claim 1, It is characterized in that Identifying a plurality of objects to be extracted in the target image, comprising: A plurality of objects to be extracted in the target image are identified according to a labeled text detection network, wherein the labels in the labeled text detection network include field name class labels and field content class labels.
5. The method according to claim 4, It is characterized in that The labels in the labeled text detection network also include background area class labels.
6. The method according to any one of claims 1 to 5, It is characterized in that The target image is at least one of a license image, a certificate image, a receipt image, and a bill image.
7. An image information output device, It is characterized in that The device comprises: an acquisition unit, a recognition unit, a matching unit and an output unit; The acquisition unit is used to acquire a target image, wherein the target image includes a plurality of field names and field contents having a structured relationship; The identification unit is used to identify a plurality of objects to be extracted in the target image, wherein the plurality of objects to be extracted include a plurality of field names and a plurality of field contents in the target image; The matching unit is used to match the multiple objects to be extracted according to the relative position constraint rule to obtain the field names and field contents that are successfully matched, specifically including: interacting the features of the object to be extracted and the features of other objects to be extracted that are adjacent to the position of the object to be extracted to obtain the interactive features of the object to be extracted; matching the multiple objects to be extracted according to the relative position constraint rule and the interactive features of the object to be extracted to obtain the field names and field contents that are successfully matched; the relative position constraint rule is used to identify the relative positions of the field names and field contents that have the structured relationship; The output unit is used to output the successfully matched field name and field content.
8. The device according to claim 7, It is characterized in that The matching unit is further specifically used for: According to the relative position constraint rule, obtaining at least one second object to be extracted that matches the first object to be extracted; Putting the first object to be extracted and the second object to be extracted into a set; If there are multiple field names in the set, determining the arrangement order of the field names in the set according to the relative positions of the field names in the set; If there are multiple field contents in the set, determining the arrangement order of the field contents in the set according to the relative positions between the field contents in the set; The first to-be-extracted objects and the second to-be-extracted objects are sorted according to the arrangement order to obtain field names and field contents that have been matched successfully.
9. The device according to claim 7, It is characterized in that Before interacting the features of the object to be extracted and the features of other objects to be extracted adjacent to the position of the object to be extracted to obtain the interactive features of the object to be extracted, the device is further used to: Extracting original features of the object to be extracted, wherein the original features include at least one of image features, position features and semantic features; If the original features of the object to be extracted include multiple types, the original features of the object to be extracted are fused to obtain the features of the object to be extracted.
10. The device according to claim 7, It is characterized in that The identification unit is specifically used for: A plurality of objects to be extracted in the target image are identified according to a labeled text detection network, wherein the labels in the labeled text detection network include field name class labels and field content class labels.
11. The device according to claim 10, It is characterized in that The labels in the labeled text detection network also include background area class labels.
12. The device according to any one of claims 7 to 11, It is characterized in that The target image is at least one of a license image, a certificate image, a receipt image, and a bill image.
13. A computer device, It is characterized in that The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1 to 6 according to the instructions in the program code.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 6.
15. A computer program product, It is characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image text recognition method and device, computer equipment and computer storage medium
CN111259889A
Image text recognition method and device, computer equipment and computer storage medium
CN111275038A