Information extraction method and device and storage medium
By matching the text information in the image with the text dictionary library and using a language model to extract it when the detailed information is not matched, the problem of insufficient accuracy of image information extraction in the prior art is solved, the extraction accuracy is improved and the processing pressure is reduced.
Patent Information
- Application Number
- CN202311734415.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, image information extraction depends on rule matching, and rules need to be updated frequently to improve accuracy. In addition, deep learning methods do not have enough understanding of short text information, and there is a problem of high computing power demand.
By obtaining text information in the image and matching it with a pre-constructed text dictionary, if the details are not matched, the image is extracted using the language model.
The extraction accuracy of the target extraction information in the image is improved, the extraction loss problem caused by incomplete text dictionary library is reduced, and the processing pressure of the language model is reduced.
Smart Images

Figure CN120164200A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision, and in particular, to an information extraction method, apparatus, and storage medium. Background Art
[0002] Image information extraction is one of the important applications in the field of computer vision. By classifying images, various objects in the images can be recognized and analyzed.
[0003] In related technologies, image information extraction is achieved by rule matching. This image information extraction method requires real-time rule updates to meet the accuracy of information extraction. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present disclosure provides an image information extraction method, apparatus, and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, an information extraction method is provided, including:
[0006] Obtain a first image and identify text information included in the first image; match the text information with a text dictionary library, where the text dictionary library is a pre-constructed text detail library including detailed information of specified class information, and the specified class information includes one or more detailed information; in response to not matching detailed information corresponding to the text information in the text dictionary library, based on a language model and the first image, extract detailed information corresponding to the specified class information of the text information in the first image.
[0007] In one implementation, the method further includes: in response to matching detailed information corresponding to the text information in the text dictionary library, use the matched detailed information as the extracted target extraction information.
[0008] In one implementation, the extracting target extraction information in the first image based on the language model and the first image includes: invoking a text input instruction for extracting target extraction information; based on the text input instruction, the first image, and the language model, obtain the target extraction information included in the first image.
[0009] In one implementation, the method further includes: in response to extracting target extraction information in the first image based on the language model, store the target extraction information in the text dictionary library.
[0010] In one implementation, before recognizing the text information included in the first image, the method further includes: determining a subject object corresponding to the target extraction information, and performing object recognition on the first image; in response to recognizing that the first image includes non-subject objects, removing the non-subject objects to obtain a second image; and using the second image as the first image for subsequent extraction of the target extraction information.
[0011] In one implementation, the target extraction information is the brand information of a product.
[0012] According to a second aspect of the embodiments of the present disclosure, there is provided an information extraction device, including:
[0013] An acquisition unit, configured to acquire a first image and recognize the text information included in the first image;
[0014] A matching unit, configured to match the text information with a text dictionary library, where the text dictionary library is a pre-constructed text information library including detailed text information of specified class information, and the specified class information includes one or more detailed information;
[0015] A processing unit, configured to, in response to not finding detailed information corresponding to the text information in the text dictionary library, extract the target extraction information corresponding to the text information from the first image based on a language model and the first image.
[0016] In one implementation, the matching unit is further configured to: in response to finding detailed information corresponding to the text information in the text dictionary library, use the found detailed information as the extracted target extraction information.
[0017] In one implementation, the processing unit extracts the target extraction information from the first image based on a language model and the first image in the following manner: invoking a text input instruction for extracting the target extraction information; and obtaining the target extraction information included in the first image based on the text input instruction, the first image, and the language model.
[0018] In one implementation, the processing unit is further configured to: in response to extracting the target extraction information from the first image based on the language model, store the target extraction information in the text dictionary library.
[0019] In one implementation, before recognizing the text information included in the first image, the acquisition unit is further configured to: determine a subject object corresponding to the target extraction information, and perform object recognition on the first image; in response to recognizing that the first image includes non-subject objects, remove the non-subject objects to obtain a second image; and use the second image as the first image for subsequent extraction of the target extraction information.
[0020] In one implementation, the target extraction information is the brand information of the product.
[0021] According to a third aspect of the embodiments of the present disclosure, an information extraction device is provided, including:
[0022] A processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the information extraction method described in the first aspect or any one of the implementations of the first aspect.
[0023] According to a fourth aspect of the embodiments of the present disclosure, a storage medium is provided, in which instructions are stored, and when the instructions in the storage medium are executed by a processor of a terminal, the terminal can perform the method described in any item of the first aspect or any one of the implementations of the first aspect.
[0024] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By matching the text information existing in the first image with the detailed information of the specified class information in the text dictionary library, if no detailed information corresponding to the text information is found, the language model is further used to extract the target extraction information from the first image to ensure that the target extraction information not included in the text dictionary library is extracted, thereby improving the extraction accuracy of the target extraction information included in the image.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0027] Figure 1 is a flowchart of an information extraction method shown according to an exemplary embodiment.
[0028] Figure 2 is a flowchart of an information extraction method shown according to an exemplary embodiment.
[0029] Figure 3 is a flowchart of a method for extracting target extraction information based on a language model shown according to an exemplary embodiment.
[0030] Figure 4 is a flowchart of another information extraction method shown according to an exemplary embodiment.
[0031] Figure 5 is a flowchart of a first image processing method shown according to an exemplary embodiment.
[0032] Figure 6 It is a schematic diagram of the process of a method for extracting product brand information shown according to an exemplary embodiment.
[0033] Figure 7 It is a schematic diagram of a first image shown according to an exemplary embodiment.
[0034] Figure 8 It is a schematic diagram of a first image of the brand information of a product containing an uncommon brand shown according to an exemplary illustration.
[0035] Figure 9 It is a block diagram of an information extraction device shown according to an exemplary embodiment.
[0036] Figure 10 It is a block diagram of a device for information extraction shown according to an exemplary embodiment. Detailed implementation manners
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present disclosure.
[0038] In the drawings, the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, of the embodiments of the present disclosure. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as a limitation of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts fall within the scope of protection of the present disclosure. The embodiments of the present disclosure will be described in detail below with reference to the drawings.
[0039] The information extraction method provided by the embodiments of the present disclosure is applied to the fields of computer vision and artificial intelligence. The information extraction method provided by the embodiments of the present disclosure is mainly used to extract specific information content in an image and can be applied to information collection in news media, scientific and technological literature, sales research, and social media.
[0040] In the embodiments of the present disclosure, the information content to be extracted in the image is referred to as target extraction information. The target extraction information is generally understood as a general term for a certain category of text information. For example, in the embodiments of the present disclosure, the target extraction information may be name information. The name information may be a product brand name. Of course, the method provided by the embodiments of the present disclosure is not limited to brand names. For example, it may also be a place name, a person's name, etc.
[0041] In related technologies, methods for extracting information by directly extracting target extraction information from the text information existing in an image or from the text information include rule-based methods, statistic-based methods, deep learning-based methods, etc. However, in related technologies, when using the method of rule-based matching to match and extract target extraction information, it depends on the construction and use of a dictionary containing the target extraction information. In this way, the requirements for the dictionary are high, and the content in the dictionary needs to be updated frequently. In the method of using deep learning to extract target extraction information, if the text information existing in the image is short, the deep learning method often has a large understanding of the text information, and there is a high computing power requirement for mining brand information from the text information. Moreover, using the method of rule-based matching or other information extraction methods often has limitations and cannot accurately extract the target extraction information to meet the needs of target extraction information extraction.
[0042] In view of this, the embodiments of the present disclosure provide an information extraction method, which combines a text dictionary library and a language model to extract target extraction information. In one implementation, the text information existing in the image is recognized, and based on a pre-constructed dictionary containing the text information corresponding to the target extraction information, the text information existing in the image is extracted. If the target extraction information is not extracted, the language model is used to extract the text information in the image again. Combining the text dictionary library and the language model to extract target extraction information can improve the accuracy of extracting the target extraction information from the text information contained in the image and reduce the situation where the target extraction information in the text information of the image is not extracted.
[0043] Figure 1 is a flowchart of an information extraction method shown according to an exemplary embodiment, as Figure 1 shown, and includes the following steps.
[0044] In step S11, a first image is obtained, and the text information included in the first image is recognized.
[0045] In the embodiments of the present disclosure, the first image is an image from which target extraction information is to be extracted. For example, it can be an image including a subject object from which target extraction information is to be extracted. Usually, the first image can be a preset image set, and this image set is an image for subsequent information collection based on the extracted target extraction information. For example, it can be an image set including a commodity as the subject object for category analysis based on the extracted brand information.
[0046] The first image includes the text information corresponding to the target extraction information to be extracted. In the embodiments of the present disclosure, the text information included in the first image is recognized through text recognition technology. For the convenience of description below, the text information recognized in the first image is directly referred to as text information.
[0047] In step S12, the text information is matched with the text dictionary library.
[0048] In the embodiments of the present disclosure, the text dictionary library is a pre-constructed text detail library including detailed information of specified class information, where the specified class information may include one or more detailed information. For example, if the specified class information is product brand information, the detailed information of the specified class information included in the text dictionary library may be brand information crawled from a specific web page through a web crawler.
[0049] In the embodiments of the present disclosure, if no detailed information corresponding to the text information is found in the text dictionary library, information extraction of the target extraction information is further performed based on the language model.
[0050] In step S13, in response to no detailed information corresponding to the text information being found in the text dictionary library, based on the language model and the first image, detailed information corresponding to the specified class information of the text information is extracted from the first image.
[0051] The language model involved in the embodiments of the present disclosure may be a pre-trained large language model. For example, a model that can understand natural language and give grammatically and semantically correct responses obtained by training with a large amount of training data.
[0052] In the embodiments of the present disclosure, in the case where no detailed information corresponding to the text information is found in the text dictionary library, the text information in the first image is secondarily extracted through the language model, reducing the occurrence of the failure to extract the target extraction information due to the lack of corresponding information content in the text dictionary library, and improving the probability of extracting the target extraction information.
[0053] In the embodiments of the present disclosure, if no corresponding detailed information is found in the text dictionary library for the text information, the language model is used to extract the target extraction information from the text information in the first image again. For example, based on the dictionary, rule matching of the product brand name is performed on the text information in Figure A, and the brand name in Figure A is not included in the dictionary, that is, it is determined that the corresponding brand name cannot be obtained using the text dictionary library. In the embodiments of the present disclosure, the language model is further used to extract the target extraction information from the text information in the first image to accurately extract the target extraction information included in the text information and improve the accuracy of obtaining the target extraction information.
[0054] It should be understood that in the embodiments of the present disclosure, the content and form of the text information included in the first image are not limited. For example, the content of the text information may be language information in multiple countries, and the form of the text information may be a deformed font obtained by processing based on ordinary text, or a handwritten font.
[0055] In the embodiments of the present disclosure, by recognizing the text information included in the first image, the text information contained in the first image is obtained. For example, the first image is scanned and recognized using Optical Character Recognition (OCR) to obtain text information that can be edited or retrieved by the electronic device. It should be understood that the text information may be multiple text segments or a table containing text information. It should be understood that using OCR to obtain the text information in the first image is only for illustrative purposes, and obtaining the text information corresponding to the first image in the embodiments of the present disclosure may also include, but is not limited to, using a trained language model to extract image features.
[0056] In the embodiments of the present disclosure, by using the detailed information in the text dictionary library as keywords to match the text content in the text information, the text content in the text information that is the same as the detailed information included in the text dictionary library is recognized, so as to match the detailed information corresponding to the text information in the text dictionary library. For example, the text dictionary library is a brand name library, and the brand name XX included in the first image is matched in the brand name library.
[0057] In the embodiments of the present disclosure, the target extraction information may include one or more detailed information. For example, if the first image contains text information corresponding to multiple target extraction information, and the content or part of the content in one or more of the text information has corresponding detailed information in the text dictionary library, then the corresponding one or more detailed information are used as the target extraction information. For example, the text information included in the first image includes brand name XX / brand name YY. The text dictionary library is a brand name library, and the brand name XX and brand name YY included in the first image are matched in the brand name library. The detailed information is brand name XX and brand name YY.
[0058] In the embodiments of the present disclosure, the extracted text information can be directly output, or post-processing operations such as removing redundant information and part-of-speech tagging can be performed on the extracted text information before output.
[0059] In the embodiments of the present disclosure, if the detailed information corresponding to the text information is matched in the text dictionary library, this detailed information can be used as the target extraction information.
[0060] It can be understood that the embodiments of the present disclosure can adopt different subsequent processing methods based on whether the detailed information corresponding to the text information is extracted in the text dictionary library.
[0061] Figure 2 is a flowchart of an information extraction method shown according to an exemplary embodiment, as Figure 2As shown, it includes the following steps: step S21, step S22, step S23a, and step S23b.
[0062] Figure 2 The steps in step S21 and S22 are the same as those in Figure 1 The steps in are the same as steps S11 and S12 in , which will not be elaborated here. You can refer to the relevant descriptions in the above embodiments. Only the differences will be described below.
[0063] In step S23a, in response to matching the detailed information corresponding to the text information in the text dictionary library, the matched detailed information is used as the extracted target extraction information.
[0064] In step S23b, in response to not matching the detailed information corresponding to the text information in the text dictionary library, based on the language model and the first image, the detailed information of the specified class corresponding to the text information is extracted from the first image.
[0065] In the embodiments of the present disclosure, extracting the target extraction information in the image based on the text dictionary library and the language model can ensure that the target extraction information included in the first image can be extracted, and improve the probability of extracting the target extraction information.
[0066] The embodiments of the present disclosure will hereinafter describe the implementation process of extracting the target extraction information based on the language model.
[0067] Figure 3 is a flowchart of a method for extracting target extraction information based on a language model shown according to an exemplary embodiment. As Figure 3 shown, it includes the following steps.
[0068] In step S31, a text input instruction for extracting the target extraction information is called.
[0069] In the embodiments of the present disclosure, the text input instruction is an instruction for guiding the language model to generate the instruction including the target extraction information.
[0070] Among them, the text input instruction can be an instruction input based on a prompt template. The text input instruction can be in the form of natural language. Among them, the text input instruction can be any form that controls the output text generated by the language model to be the target extraction information.
[0071] In one example, the target extraction information is brand information, and the text input instruction can be: Please extract the brand information included in the image.
[0072] In step S32, based on the text input instruction, the first image, and the language model, the target extraction information included in the first image is obtained.
[0073] In the embodiments of the present disclosure, a text input instruction and a first image can be input into a language model as input content of the language model, and target extraction information included in the first image can be determined according to the output result of the language model.
[0074] Among them, in the embodiments of the present disclosure, all or part of the text information included in the first image can be input into the language model as the input content of the language model to obtain the target extraction information included in the first image. For example, if there are multiple groups of text information including target extraction information in the first image. Among them, if detailed information matching exists for part of the groups of text information in the text dictionary library, the text information for which no detailed information is matched in the text dictionary library can be used as the input content for input into the language model.
[0075] In the embodiments of the present disclosure, after initially matching the target extraction information through the text dictionary library and then extracting the target extraction information based on the language model, the processing pressure of the language model can be reduced and the extraction cost can be lowered.
[0076] In the embodiments of the present disclosure, the target extraction information extracted through the first image and the language model can be used as new detailed information to be input into the text dictionary library to update the text dictionary library, improve the richness of the text dictionary library, and increase the accuracy of obtaining the target extraction information by matching the text information using the text dictionary library.
[0077] Figure 4 is a flowchart of another information extraction method shown according to an exemplary embodiment, as Figure 4 shown, including the following steps: step S41, step S42, step S43, and step S44.
[0078] Among them, Figure 4 the steps in steps S41, S42, and S43 in Figure 1 are the same as the steps S11, S12, and step S13 in
[0079] and will not be elaborated here. Reference can be made to the relevant descriptions of the above embodiments.
[0080] In the embodiments of the present disclosure, if some groups of text information included in the first image match the detailed information in the text dictionary library, while other groups of text information do not match the detailed information in the text dictionary library. When the input content of the language model is the first image and the text input instruction, the target extraction information obtained by the language model includes all the text information in the first image. It is necessary to filter the target extraction information output by the language model, remove the content of the target extraction information that already exists in the text dictionary library, and then store the filtered target extraction information in the text dictionary library to ensure that the stored text information is not yet included in the text dictionary library and remove redundant data.
[0081] In the embodiments of the present disclosure, the first image for extracting target extraction information can be the original image input by the user, or the second image obtained by preprocessing the original image input by the user.
[0082] In the embodiments of the present disclosure, to ensure the accuracy of the extracted target extraction information, after obtaining the first image, the first image can be preprocessed to exclude the interference and processing pressure caused by useless information in the first image, obtain the second image, and use the second image as the extraction object for subsequent target extraction information to perform matching in the text dictionary library and / or information extraction of the language model.
[0083] Figure 5 is a flowchart of a first image processing method shown according to an exemplary embodiment, as Figure 5 shown, including the following steps.
[0084] In step S51, determine the subject object corresponding to the target extraction information, and perform object recognition on the first image.
[0085] In the embodiments of the present disclosure, the method for determining the corresponding subject object according to the target extraction information can be to perform matching based on preset specific rules and patterns, train a classifier or regression model based on existing data, and create a knowledge graph, etc., to determine the subject object and non-subject object in the first image. Among them, the non-subject object is determined based on the subject object in the first image, and there is no common part between the subject object and the non-subject object. For example, if the subject object corresponding to the target extraction information is the target commodity, the non-subject object can be other commodities other than the target commodity or other items other than commodities.
[0086] In one example, the subject object is a commodity of category M, and the non-subject object can be other commodities of the same category but different from category M.
[0087] In step S52, in response to recognizing that the first image includes a non-subject object, remove the non-subject object to obtain the second image.
[0088] In the embodiments of the present disclosure, if the non-subject object is other content than a commodity, the second image only contains the subject object; if the non-subject object is other articles than a commodity, and the other articles than the commodity in the first image are gifts, the content in the second image is the content obtained by erasing the gift image from the first image.
[0089] In the embodiments of the present disclosure, the non-subject object can be separated by an image segmentation algorithm through an image segmentation method, and then deleted or masked from the image.
[0090] In step S53, the second image is used as the first image for extracting subsequent target information.
[0091] In the embodiments of the present disclosure, the second image obtained after processing the first image is used as the new first image, and text information is recognized to avoid waste of operations and resources caused by recognizing and extracting text information in the non-subject object.
[0092] The implementation process of using the second image as the first image for subsequent target information extraction in the embodiments of the present disclosure can refer to the information extraction implementation process involved in any one of the above embodiments Figures 1 to 4 and will not be elaborated herein.
[0093] In the embodiments of the present disclosure, taking the target extraction information as the brand information of a commodity as an example, the implementation process involved in the above embodiments will be described.
[0094] Figure 6 It is a schematic diagram of the process of a commodity brand information extraction method shown according to an exemplary embodiment.
[0095] In Figure 6Among them, the target extraction information is the brand information of the commodity. The first image can be called various material images including the commodity. The text dictionary library is the brand dictionary, and the specified class information is the brand name. By using a crawler robot to obtain the brand information contained in website C as detailed information, and pre-constructing the brand dictionary. When extracting brand information, input the material image containing commodity information, and recognize multiple groups of text information (the first group of commodity brand information and the second group of commodity brand information) in the material through OCR. Match each group of text information with the brand names contained in the brand dictionary respectively. If it is detected that there is the first group of commodity brand information in the first group of text information that matches the brand names contained in the brand dictionary, then save the matched first group of commodity brand information as one of the output contents. Temporarily store the second group of brand information of the commodity brand information that is not matched with the brand names contained in the brand dictionary, and combine the second group of brand information with the text input instruction to obtain the input content of the language model and input it into the language model for brand information extraction. Among them, the content of the text input instruction can be, for example, "Does the image contain brand information? If not, please answer no. If so, please output the brand information". Input the input content into the language model, and through the semantic understanding of the language model, extract the second group of commodity brand information. In the embodiment of the present disclosure, after extracting the first group of commodity brand information and the second group of commodity brand information, output the first group of commodity brand information and the second group of commodity brand information.
[0096] Through the information extraction method provided by the embodiment of the present disclosure, it is possible to apply to the extraction of brand information in various forms and improve the accuracy of information extraction. Figure 7 is a schematic diagram of the first image shown according to an exemplary embodiment. In Figure 7 Among them, brand A is the brand information of the commodity, and brand A is English information. The area where brand A is located is the entity area. When the corresponding English information of brand A does not exist in the text dictionary library, it is impossible to match the English information of brand A through the text dictionary library to determine the brand information of the matched commodity. However, by inputting the English information of brand A and the text input instruction into the language model, it is possible to determine that the brand information of the commodity is brand A.
[0097] Through the information extraction method provided by the embodiment of the present disclosure, it is possible to extract the brand information of commodities with uncommon brands. Figure 8 is a schematic diagram of the first image containing the brand information of a commodity with an uncommon brand shown according to an exemplary illustration. In Figure 8In it, the area where "Brand B products" and the product pictures are located is the entity area, and other areas are non-entity areas. However, the brand information of the products of Brand B is not included in the text dictionary library, resulting in the inability to match Brand B through the text dictionary library. At this time, the image in the entity area is combined with the content of the text input instruction "Does the image contain brand information? If not, please answer no. If so, please output the brand information" to obtain the input content of the language model. The input content of the language model is input into the language model, and the language model processes the image in the entity area and the content of the text input instruction to obtain the output brand information of the product, that is, the words "Brand B".
[0098] In the brand information extraction method provided by the embodiments of the present disclosure, the image can be preprocessed to eliminate redundant information. For example Figure 7 the sticker in Figure 2 is an area including a gift image and text information, and the Figure 2 text information in also contains brand information. However, since the Figure 2 is a gift and not a product, the content in the Figure 2 is interference information or redundant information. The area of the Figure 2 sticker can be marked through manual or automated rules, and this area can be set as a non-entity area to avoid obtaining redundant information such as the brand information of non-products later, which affects the quality of extracting the brand information of products.
[0099] In the embodiments of the present disclosure, by matching the text dictionary library with the text information in the first image, and detecting the target extraction information contained in the first image that is not matched again through the language model, the problem of missing target extraction information caused by the imperfect detailed information in the text dictionary library can be reduced. And by matching with the text dictionary library before using the language model, the problem of high resource cost and computing cost required for extracting all text information through the language model can be reduced. Moreover, based on the text dictionary library and the language model, the information extraction of the text information in the first image is improved, and the accuracy of extracting the target extraction information is improved.
[0100] Based on the same concept, the embodiments of the present disclosure also provide an information extraction device.
[0101] It is understandable that, in order to implement the above functions, the information extraction device provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for executing each function. Combining the units and algorithm steps of the various examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0102] Figure 9 It is a block diagram of an information extraction device 100 shown according to an exemplary embodiment. Referring to Figure 9 this device includes an acquisition unit 101, a matching unit 102, and a processing unit 103.
[0103] The acquisition unit 101 is used to acquire a first image and identify the text information included in the first image.
[0104] The matching unit 102 is used to match the text information with a text dictionary library, where the text dictionary library is a pre-constructed text information library including detailed text information of specified class information, and the specified class information includes one or more detailed information.
[0105] The processing unit 103 is used to, in response to not matching the detailed information corresponding to the text information in the text dictionary library, extract the detailed information corresponding to the specified class information of the text information in the first image based on the language model and the first image.
[0106] In one embodiment, the matching unit 102 is further used for:
[0107] In response to matching the detailed information corresponding to the text information in the text dictionary library, taking the matched detailed information as the extracted target extraction information.
[0108] In one embodiment, the processing unit 103 extracts the target extraction information in the first image based on the language model and the first image in the following manner:
[0109] Call the text input instruction for extracting the target extraction information; based on the text input instruction, the first image, and the language model, obtain the target extraction information included in the first image.
[0110] In one embodiment, the processing unit 103 is further used for:
[0111] In response to extracting the target extraction information in the first image based on the language model, storing the target extraction information in the text dictionary library.
[0112] In one embodiment, before identifying the text information included in the first image, the obtaining unit 101 is further configured to:
[0113] Determine the subject object corresponding to the target extraction information, and perform object recognition on the first image; in response to identifying that the first image includes non-subject objects, remove the non-subject objects to obtain a second image; and use the second image as the first image for subsequent target extraction information.
[0114] In one embodiment, the target extraction information is the brand information of a commodity.
[0115] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0116] Figure 10 It is a block diagram of a device 200 for information extraction shown according to an exemplary embodiment. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0117] Referring to Figure 10 , the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0118] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0119] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, and the like. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0120] The power component 206 provides power for various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.
[0121] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0122] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0123] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0124] The sensor assembly 214 includes one or more sensors for providing an assessment of the status of the device 200 in various aspects. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0125] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0126] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0127] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 204 including instructions, is also provided. The above instructions can be executed by the processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0128] It is understood that in the present disclosure, "a plurality of" means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0129] It can be further understood that the terms "first", "second", etc. are used to describe various information, but such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other and do not represent a specific order or degree of importance. In fact, expressions such as "first" and "second" can be used interchangeably completely. For example, without departing from the scope of the present disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.
[0130] It can be further understood that unless otherwise specified, "connection" includes direct connection without other components between the two, and also includes indirect connection with other elements between the two.
[0131] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be beneficial.
[0132] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0133] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An information extraction method, characterized in that, Including: Obtain a first image and identify the text information included in the first image; Match the text information with a text dictionary library, which is a pre-constructed text detail library including detailed information of specified class information, and the specified class information includes one or more detailed information; In response to not matching the detailed information corresponding to the text information in the text dictionary library, based on the language model and the first image, extract the detailed information corresponding to the specified class information of the text information in the first image.
2. The method according to claim 1, characterized in that, The method further includes: In response to matching the detailed information corresponding to the text information in the text dictionary library, use the matched detailed information as the target extraction information extracted.
3. The method according to claim 1, characterized in that, The extracting the target extraction information in the first image based on the language model and the first image includes: Invoke a text input instruction for extracting the target extraction information; Based on the text input instruction, the first image and the language model, obtain the target extraction information included in the first image.
4. The method according to any one of claims 1 - 3, characterized in that, The method further includes: In response to extracting the target extraction information in the first image based on the language model, store the target extraction information in the text dictionary library.
5. The method according to claim 1, characterized in that, Before identifying the text information included in the first image, the method further includes: Determine the main object corresponding to the target extraction information and perform object recognition on the first image; In response to identifying that the first image includes non-main objects, remove the non-main objects to obtain a second image; Use the second image as the first image for subsequent extraction of the target extraction information.
6. The method according to claim 2, characterized in that, The target extraction information is the brand information of the commodity.
7. An information extraction device, characterized in that, Including: An acquisition unit for obtaining a first image and identifying the text information included in the first image; A matching unit for matching the text information with a text dictionary library, which is a pre-constructed text information library including detailed text information of specified class information, and the specified class information includes one or more detailed information; A processing unit for, in response to not matching the detailed information corresponding to the text information in the text dictionary library, based on the language model and the first image, extracting the detailed information corresponding to the specified class information of the text information in the first image.
8. The device according to claim 7, characterized in that, The matching unit is further used for: In response to matching the detailed information corresponding to the text information in the text dictionary library, using the matched detailed information as the target extraction information extracted.
9. The device according to claim 7, characterized in that, The processing unit extracts the target extraction information in the first image based on the language model and the first image in the following manner: Invoke a text input instruction for extracting the target extraction information; Based on the text input instruction, the first image and the language model, obtain the target extraction information included in the first image.
10. The device according to any one of claims 7 - 9, characterized in that, The processing unit is further used for: In response to extracting the target extraction information in the first image based on the language model, storing the target extraction information in the text dictionary library.
11. The device according to claim 7, characterized in that, Before identifying the text information included in the first image, the acquisition unit is further used for: Determine the subject object corresponding to the target extraction information, and perform object recognition on the first image; In response to recognizing that the first image includes non-subject objects, remove the non-subject objects to obtain a second image; Use the second image as the first image for subsequent extraction of the target extraction information.
12. The device according to claim 8, characterized in that, The target extraction information is the brand information of the commodity.
13. An information extraction device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the method described in any one of claims 1 to 6.
14. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions in the storage medium are executed by the processor of the terminal, the terminal is enabled to execute the method described in any one of claims 1 to 6.
Citation Information
Cited By
OCR (Optical Character Recognition) text recognition method and device based on hot word perception, equipment and medium
CN121074918A