Information extraction method, device and medium
By using positioning convolution kernels of different scales to process historical positioning information in information extraction, the problem of low accuracy of feature information extraction caused by inconsistent size of information units in the image is solved, and higher accuracy of information extraction is achieved.
Patent Information
- Application Number
- CN202111275885.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the information extraction scenario, the size of the information units in the image is inconsistent, resulting in low accuracy of feature information extraction.
By obtaining the historical positioning information of the information unit that has been extracted in the target image, and using at least two positioning convolutions of different scales to convolution process the historical positioning information, the positioning reference information is obtained, and the target positioning information of the information unit that has not been extracted in the feature information is determined and its characteristic information is extracted.
The accuracy of information extraction is improved, and the positioning convolution kernels of different scales focus on information units of different sizes is ensured, thus improving the accuracy of feature information extraction.
Smart Images

Figure CN114155295B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and artificial intelligence technology, and more specifically, to an information extraction method, device and medium. Background Art
[0002] In information extraction scenarios, such as information extraction scenarios in images (for example, extracting formulas or text in images), information units in images usually have inconsistent sizes, which will hinder the extraction of feature information of the information units, thereby resulting in low accuracy in the extraction of feature information of the information units.
[0003] Based on this, how to improve the accuracy of information extraction is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present application provide an information extraction method, apparatus, computer program product or computer program, and computer-readable medium, which can improve the accuracy of information extraction at least to a certain extent.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of an embodiment of the present application, an information extraction method is provided, comprising: obtaining historical positioning information of a first information unit from a target image from which feature information extraction has been completed, the target image comprising at least one information unit; performing convolution processing on the historical positioning information using at least two positioning convolution kernels of different scales, respectively, to obtain at least two positioning reference information; determining target positioning information of a second information unit from which feature information extraction has not been completed in the target image based on the at least two positioning reference information; and extracting target feature information of the second information unit based on the target positioning information.
[0007] According to one aspect of an embodiment of the present application, an information extraction device is provided, comprising: an acquisition unit, used to acquire historical positioning information of a first information unit from a target image from which feature information extraction has been completed, wherein the target image includes at least one information unit; a convolution unit, used to perform convolution processing on the historical positioning information through at least two positioning convolution kernels of different scales, respectively, to obtain at least two positioning reference information; a determination unit, used to determine target positioning information of a second information unit from which feature information extraction has not been completed in the target image based on the at least two positioning reference information; and an extraction unit, used to extract target feature information of the second information unit based on the target positioning information.
[0008] In some embodiments of the present application, based on the aforementioned scheme, the acquisition unit is configured to: obtain the historical positioning information of each first information unit that has completed feature information extraction in the target image; calculate the historical positioning information of each first information unit that has completed feature information extraction to obtain the historical positioning information.
[0009] In some embodiments of the present application, based on the aforementioned scheme, the convolution unit is configured to: perform convolution processing on the historical positioning information through the first positioning convolution kernel in the information extraction model to obtain first positioning reference information; perform convolution processing on the historical positioning information through the second positioning convolution kernel in the information extraction model to obtain second positioning reference information, and the scale of the second positioning convolution kernel is larger than the scale of the first positioning convolution kernel.
[0010] In some embodiments of the present application, based on the aforementioned scheme, the determination unit is configured to: before determining the target positioning information of the second information unit in the target image for which feature information extraction has not been completed based on the at least two positioning reference information, determine the second information unit from the information units in the target image for which feature information extraction has not been completed according to a preset arrangement direction of the at least one information unit in the target image.
[0011] In some embodiments of the present application, based on the aforementioned scheme, the determination unit is further configured to: obtain hidden state information of the information unit that is closest to the second information unit when extracting feature information, and obtain encoded feature data for the target image; aggregate the at least two positioning reference information, the hidden state information, and the encoded feature data to obtain the target positioning information of the second information unit.
[0012] In some embodiments of the present application, based on the aforementioned scheme, the extraction unit is configured to: obtain encoded feature data for the target image; based on the target positioning information, decode the encoded feature data through a target decoder model in the information extraction model to obtain target feature information of the second information unit.
[0013] In some embodiments of the present application, based on the aforementioned scheme, the device also includes: a training unit, which is used to obtain a model to be trained before obtaining historical positioning information of a first information unit from which feature information extraction has been completed in a target image, the model to be trained comprising an encoder model and at least two decoder models, the encoder model being used to encode an image to obtain encoded feature data, and the decoder model being used to decode the encoded feature data to obtain feature information of each information unit in the image; obtaining a sample image, and training the model to be trained using the sample image to obtain an information extraction model.
[0014] In some embodiments of the present application, based on the aforementioned scheme, the encoder model includes a densely connected convolutional network model.
[0015] In some embodiments of the present application, based on the aforementioned scheme, the training unit is configured to: encode the sample image through the encoder model to obtain sample encoding feature data; decode the sample encoding feature data through the at least two decoder models respectively to obtain at least two groups of sample feature information, wherein each group of sample feature information includes feature information for each information unit in the sample image; trigger each of the at least two decoder models to learn the sample feature information decoded by other decoder models except itself; determine a decoder model as a target decoder model among the at least two decoder models to obtain the information extraction model composed of the encoder model and the target decoder model.
[0016] In some embodiments of the present application, based on the aforementioned scheme, the training unit is configured as: based on the sample coding feature data, determining the sample positioning information of each information unit in the sample image according to different arrangement directions of the information units in the sample image through the at least two decoder models; based on the sample positioning information corresponding to the at least two decoder models, respectively decode the sample coding feature data to obtain the at least two groups of sample feature information.
[0017] In some embodiments of the present application, based on the aforementioned scheme, the information unit includes a character unit, at least one information unit in the target image constitutes one or more formulas containing the character unit, and the at least two positioning reference information are used to respectively focus on the positioning information of character units of different sizes.
[0018] In some embodiments of the present application, based on the aforementioned scheme, the device also includes: an editing unit, which is used to obtain target feature information corresponding to each character unit in the target image after extracting the target feature information of the second information unit based on the target positioning information; based on the target feature information, edit one or more formulas in the target image to a formula editing area.
[0019] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the information extraction method described in the above embodiments.
[0020] According to one aspect of an embodiment of the present application, an information extraction device is also provided, characterized in that it includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing the information extraction method described in the above embodiments.
[0021] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one program code is stored. The at least one program code is loaded and executed by a processor to implement the operations performed by the information extraction method described in the above embodiment.
[0022] In the technical solutions provided in some embodiments of the present application, at least two positioning reference information can be obtained by convolving the historical positioning information of the first information unit whose feature information has been extracted in the target image through at least two positioning convolution kernels of different scales, and then the target positioning information of the second information unit whose feature information has not been extracted in the target image can be determined based on the at least two positioning reference information, and finally the target feature information of the second information unit can be extracted based on the target positioning information. Since the historical positioning information is convolved with at least two positioning convolution kernels of different scales when determining the target positioning information of the second information unit, the positioning convolution kernels of different scales can be used to focus on information units of different sizes in the target image, so the obtained at least two positioning reference information also include the focus information for information units of different sizes, so that the target positioning information with higher accuracy for the second information unit can be determined based on the at least two positioning reference information, thereby improving the accuracy of the target feature information of the second information unit determined by the target positioning information.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0025] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied;
[0026] Figure 2 A flow chart of an information extraction method according to an embodiment of the present application is shown;
[0027] Figure 3 A detailed flow chart of obtaining historical positioning information of a first information unit from which feature information extraction has been completed in a target image according to an embodiment of the present application is shown;
[0028] Figure 4 A detailed flow chart of performing convolution processing on the historical positioning information according to an embodiment of the present application is shown;
[0029] Figure 5 A schematic diagram showing a method of determining the second information unit from information units in the target image where feature information extraction has not been completed according to an embodiment of the present application is shown;
[0030] Figure 6 A detailed flow chart of determining target positioning information of a second information unit in the target image for which feature information extraction has not been completed according to an embodiment of the present application is shown;
[0031] Figure 7 A schematic diagram showing a framework for determining target positioning information of a second information unit in the target image for which feature information extraction has not been completed according to an embodiment of the present application;
[0032] Figure 8 A detailed flow chart of extracting target feature information of the second information unit according to an embodiment of the present application is shown;
[0033] Fig. 9 A flowchart of a method before acquiring historical positioning information of a first information unit whose feature information has been extracted in a target image according to an embodiment of the present application is shown;
[0034] Fig.10 A detailed flow chart of training the model to be trained by using the sample image according to an embodiment of the present application is shown;
[0035] Fig.11 A detailed flow chart of respectively decoding the sample encoding feature data by the at least two decoder models according to an embodiment of the present application is shown;
[0036] Fig.12 A schematic diagram of a framework of a training information extraction model according to an embodiment of the present application is shown;
[0037] Fig.13 A flowchart of a method after extracting target feature information of the second information unit based on the target positioning information according to an embodiment of the present application is shown;
[0038] Fig.14 A block diagram of an information extraction device according to an embodiment of the present application is shown;
[0039] Fig.15 A block diagram of an information extraction device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0041] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0042] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0043] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0044] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the objects used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those shown or described.
[0046] The embodiments in this application involve technologies related to artificial intelligence, that is, fully automated processing of data (such as image data) is achieved through artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0047] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.
[0048] like Figure 1 As shown, the system architecture may include terminal devices (such as Figure 1 The embodiment of the present invention is a network 104 and a server 105. The network 104 is a medium for providing a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as a wired communication link, a wireless communication link, etc.
[0049] In one embodiment of the present application, when a user needs to identify feature information reflected by at least one information unit in a target image, the user can send the target image including at least one information unit to the server 105 through a terminal device. After acquiring the target image, the server 105 extracts the feature information reflected by at least one information unit in the target image. The scheme may be: acquiring historical positioning information of a first information unit in the target image from which feature information extraction has been completed, wherein the target image includes at least one information unit; performing convolution processing on the historical positioning information respectively through at least two positioning convolution kernels of different scales to obtain at least two positioning reference information; determining the target positioning information of a second information unit in the target image from which feature information extraction has not been completed based on the at least two positioning reference information; and extracting the target feature information of the second information unit based on the target positioning information.
[0050] In this embodiment, when determining the target positioning information of the second information unit, the historical positioning information is convolved respectively by at least two positioning convolution kernels of different scales. Positioning convolution kernels of different scales can be used to focus on information units of different sizes in the target image. Therefore, the at least two positioning reference information obtained also include attention information for information units of different sizes, so that based on the at least two positioning reference information, target positioning information with higher accuracy for the second information unit can be determined, thereby improving the accuracy of the target feature information of the second information unit determined by the target positioning information.
[0051] It should be noted that the information extraction method provided in the embodiment of the present application can be executed by the server 105, and accordingly, the information extraction device is generally arranged in the server 105. However, in other embodiments of the present application, the terminal device can also have similar functions as the server, so as to execute the information extraction scheme provided in the embodiment of the present application.
[0052] It should also be noted that Figure 1 The number of terminal devices, networks and servers in the description is only for illustration. According to the implementation requirements, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0053] It should be explained that cloud computing as described above is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" are infinitely expandable in the eyes of users, and can be obtained at any time, used on demand, and expanded at any time. By establishing a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform, various types of virtual resources are deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0054] The implementation details of the technical solution of the embodiment of the present application are described in detail below:
[0055] Figure 2 A flowchart of an information extraction method according to an embodiment of the present application is shown. The information extraction method can be executed by a device having a computing processing function, such as Figure 1 , or may be performed by the server 105 shown in Figure 1 Refer to the terminal device shown in Figure 2 As shown, the information extraction method at least includes steps 220 to 280, which are described in detail as follows:
[0056] In step 220, historical positioning information of a first information unit from which feature information extraction has been completed in a target image is obtained, wherein the target image includes at least one information unit.
[0057] In the present application, the target image may be obtained by capturing a page area containing a target object in the interface, or may be directly obtained locally, and the target image includes at least one information unit.
[0058] In the present application, the proposed information extraction scheme can be applied to information recognition scenarios of target objects in images, such as formula recognition scenarios in images, text recognition scenarios in images, and certain specific pattern recognition scenarios in images. Furthermore, the target object in the image can be composed of at least one information unit, such as a formula or text in an image can be composed of at least one character unit, and certain specific patterns in an image can be composed of at least one graphic unit.
[0059] It should be noted that the feature information proposed in the present application may include the shape feature information of the information unit itself, may include the position feature information of the information unit (for example, the relative position relationship with other information units), and may also include the shape feature information and position feature information of the information unit itself. For example, taking the formula recognition scenario in the image as an example, the information unit may be a character unit, and at least one information unit in the target image constitutes one or more formulas containing the character unit. It can be understood that the feature information of the character unit in the formula may include the shape feature information of each character unit and / or the feature information of the relative position relationship between the character units.
[0060] It can be understood that, in the present application, each information unit in the target image corresponds to a positioning information in the target image, and before extracting feature information from the information unit, the positioning information of the information unit needs to be determined.
[0061] It should be noted that, in the process of extracting the characteristic information of the information unit in the target image, the positioning information of an information unit may be determined first, and the characteristic information of the information unit defined by the positioning information may be extracted. Then, the positioning information of the next information unit may be determined, and the characteristic information of the next information unit may be extracted. In this way, the characteristic information of the information units may be extracted step by step until the characteristic information of all the information units in the target image is extracted.
[0062] In order to enable those skilled in the art to better understand the present application, the following continues to explain using the formula recognition scenario in an image as an example.
[0063] For example, for the character "B" and the character "-" in the formula "A+BC", the positioning information of the character "B" is determined first, and the feature information of the character "B" is extracted based on the positioning information of the character "B". Then, the positioning information of the character "-" is determined, and the feature information of the character "-" is extracted based on the positioning information of the character "-".
[0064] In such Figure 2 In one embodiment of step 220, the historical positioning information of the first information unit from which feature information extraction has been completed in the target image is obtained, which can be performed as follows: Figure 3 Follow the steps shown.
[0065] See also Figure 3 , shows a detailed flow chart of obtaining historical positioning information of an information unit in a target image that has completed feature information extraction according to an embodiment of the present application. Specifically, it includes steps 221 to 222:
[0066] Step 221, obtaining historical positioning information of each first information unit whose feature information has been extracted in the target image.
[0067] Step 222: Calculate the historical positioning information of each first information unit from which feature information has been extracted to obtain the historical positioning information.
[0068] In this embodiment, the formula recognition scenario in the image is continued as an example for explanation. For example, for the formula "A+BC", if the feature information extraction for the character "A", the character "+", and the character "B" has been completed in history, the historical positioning information of the character "A", the character "+", and the character "B" is obtained, and the historical positioning information of the character "A", the character "+", and the character "B" is calculated (for example, added) to obtain the historical positioning information.
[0069] It should be noted that the positioning information mentioned in this application can essentially be represented by a matrix information.
[0070] Continue to refer to Figure 2 In step 240, the historical positioning information is convolved using at least two positioning convolution kernels of different scales to obtain at least two positioning reference information.
[0071] In the present application, the information extraction model may include at least two localization convolution kernels of different scales, and the localization convolution kernels of different scales can be used to focus on information units of different sizes in the target image.
[0072] It can be understood that in the present application, if the information unit includes a character unit, and at least one information unit in the target image constitutes one or more formulas containing the character unit, then the historical positioning information is convolved with at least two positioning convolution kernels of different scales to obtain at least two positioning reference information that can be used to focus on the positioning information of character units of different sizes.
[0073] In such Figure 2 In one embodiment of step 240, at least two positioning convolution kernels of different scales are used to convolve the historical positioning information to obtain at least two positioning reference information. Figure 4 Follow the steps shown.
[0074] See also Figure 4 , shows a detailed flow chart of performing convolution processing on the historical positioning information according to an embodiment of the present application. Specifically, it includes steps 241 to 242:
[0075] Step 241: Perform convolution processing on the historical positioning information through the first positioning convolution kernel in the information extraction model to obtain first positioning reference information.
[0076] Step 242: Convolution processing is performed on the historical positioning information through a second positioning convolution kernel in the information extraction model to obtain second positioning reference information, where the scale of the second positioning convolution kernel is larger than the scale of the first positioning convolution kernel.
[0077] Further, for example, in this embodiment, the scale size of the first positioning convolution kernel may be 5×5, and the scale size of the second positioning convolution kernel may be 11×11.
[0078] It can be seen that in this embodiment, the number of positioning convolution kernels is two, but in other embodiments, the number of positioning convolution kernels can also be three or four. Here, this application does not make a specific limitation on the number of positioning convolution kernels.
[0079] In the present application, the historical positioning information is convolved by positioning convolution kernels of different scales, which can focus on information units of different sizes in the target image, so that more accurate positioning information can be determined for the second information unit (i.e., a target information unit) in the subsequent process.
[0080] Continue to refer to Figure 2 In step 260, based on the at least two positioning reference information, the target positioning information of the second information unit in the target image whose feature information extraction has not been completed is determined.
[0081] In such Figure 2 In one embodiment of step 260, before determining the target positioning information of the second information unit in the target image for which feature information extraction has not been completed based on the at least two positioning reference information, the following steps may be performed:
[0082] According to a preset arrangement direction of the at least one information unit in the target image, the second information unit is determined from information units in the target image for which feature information extraction has not been completed.
[0083] In the present application, the information units in the target image are arranged in a fixed order. In order to enable those skilled in the art to better understand the present application, the following continues to take the formula recognition scene in the image as an example. Figure 5 Provide explanation.
[0084] See also Figure 5 , showing a schematic diagram of determining the second information unit from information units in which feature information extraction has not been completed in the target image according to an embodiment of the present application.
[0085] like Figure 5As shown, for the formula "A+BC" 500, the character "+" is arranged on the right side of the character "A", the character "B" is arranged on the right side of the character "+", the character "-" is arranged on the right side of the character "B", and the character "C" is arranged on the right side of the character "-".
[0086] Furthermore, the preset arrangement direction of the at least one information unit in the target image may be a preset arrangement direction of "A"→"+"→"B"→"-"→"C", or a preset arrangement direction of "C"→"-"→"B"→"+"→"A".
[0087] Furthermore, in the present embodiment, the second information unit can be determined from the information units from which feature information extraction has not been completed in the target image according to the preset arrangement direction of "A" → "+" → "B" → "-" → "C". If the character "A" and the character "+" in the formula "A+BC" have completed feature information extraction, then the character "B", the character "-" and the character "C" are information units from which feature information extraction has not been completed. It can be understood that according to the preset arrangement direction, the character "B" can be determined as the second information unit.
[0088] In such Figure 2 In one embodiment of step 260, based on the at least two positioning reference information, the target positioning information of the second information unit in the target image for which feature information extraction has not been completed can be determined as follows: Figure 6 Follow the steps shown.
[0089] See also Figure 6 , shows a detailed flow chart of determining the target positioning information of the second information unit in the target image for which feature information extraction has not been completed according to an embodiment of the present application. Specifically, it includes steps 261 to 262:
[0090] Step 261, obtaining hidden state information of the information unit that is closest to the second information unit during feature information extraction, and obtaining encoded feature data for the target image.
[0091] Step 263: Aggregate the at least two positioning reference information, the hidden state information, and the encoding feature data to obtain the target positioning information of the second information unit.
[0092] In order to make those skilled in the art better understand the present application, Figure 7 This is explained with a specific example.
[0093] See also Figure 7 , shows a schematic diagram of a framework for determining target positioning information of a second information unit in the target image for which feature information extraction has not been completed according to an embodiment of the present application.
[0094] like Figure 7 As shown, formula 701 represents the historical positioning information of the character "A" and the character "+" in the formula "A+BC". By calculating the positioning information of the character "A" and the character "+", the historical positioning information represented by formula 702 is obtained. The historical positioning information represented by formula 702 is convolved by a 5×5 positioning convolution kernel 703 and an 11×11 positioning convolution kernel 704 respectively to obtain a first positioning reference information 706 and a second positioning reference information 707. At the same time, the hidden state information 705 of the information unit (i.e., the character "+") that is closest to the second information unit (i.e., the character "B") during feature information extraction is obtained, and the encoded feature data 708 for the target image is obtained. Finally, the first positioning reference information 706, the second positioning reference information 707, the hidden state information 705 and the encoded feature data 708 are aggregated to obtain the target positioning information 709 of the second information unit (i.e., the character "B").
[0095] In this embodiment, by aggregating the at least two positioning reference information, the hidden state information, and the encoded feature data, target positioning information of the second information unit with higher accuracy can be obtained, thereby improving the accuracy of the target feature information of the second information unit determined by the target positioning information.
[0096] Continue to refer to Figure 2 In step 280, target feature information of the second information unit is extracted based on the target positioning information.
[0097] In such Figure 2 In one embodiment of step 280, based on the target positioning information, the target feature information of the second information unit is extracted as follows: Figure 8 Follow the steps shown.
[0098] See also Figure 8 , shows a detailed flow chart of extracting target feature information of the second information unit according to an embodiment of the present application. Specifically including steps 281 to 282:
[0099] Step 281, obtaining encoding feature data for the target image.
[0100] Step 282: Based on the target positioning information, the encoded feature data is decoded by a target decoder model in the information extraction model to obtain target feature information of the second information unit.
[0101] It should be noted that the information extraction model includes an encoder model and a target decoder model, the encoder model is used to encode the target image to obtain encoded feature data, and the target decoder model is used to decode the encoded feature data to obtain feature information of each information unit in the image, wherein the positioning information of each information unit in the target image is determined based on the target decoder model.
[0102] In this application, the proposed information extraction scheme is mainly implemented based on a pre-trained information extraction model, and the following will be explained in detail based on the information extraction model.
[0103] In one embodiment of the present application, before acquiring the historical positioning information of the first information unit from which feature information extraction has been completed in the target image, the following may be performed: Fig. 9 Steps shown.
[0104] See also Fig. 9 , shows a flowchart of a method before acquiring historical positioning information of a first information unit from which feature information extraction has been completed in a target image according to an embodiment of the present application. Specifically, it includes steps 200 to 210:
[0105] Step 200, obtain a model to be trained, wherein the model to be trained includes an encoder model and at least two decoder models, the encoder model is used to encode the image to obtain encoded feature data, and the decoder model is used to decode the encoded feature data to obtain feature information of each information unit in the image.
[0106] Step 210: Acquire a sample image, and train the model to be trained using the sample image to obtain an information extraction model.
[0107] In the present application, the encoder model can be searched in advance through the network structure, and then the model to be trained is constructed based on the encoder model and at least two decoder models. Those skilled in the art can understand that the encoder model can essentially belong to the network structure model.
[0108] In this application, Neural Architecture Search (NAS) is an effective tool for generating and optimizing network structures. When the length and structure of the network are uncertain, a recurrent neural network is used as a controller to generate fields of the network structure to construct a sub-neural network. The accuracy after training the sub-network is used as the controller feedback signal, and the controller is updated by calculating the policy gradient, and the iterative cycle is repeated continuously. In the next iteration, the controller will have a higher probability of proposing a high-accuracy network structure. Based on this, the encoder model is obtained by network structure search, which has the advantage of obtaining a better encoder model, so that the constructed model to be trained has accurate learning ability.
[0109] In one embodiment of the present application, the encoder model may include any one of a densely connected convolutional network model (DenseNet), a MobileNetV2 model, and an Xception model.
[0110] In one embodiment of the present application, the decoder model may include any one of a GRU model, an LSTM model, and a Transformer model.
[0111] In one embodiment of the present application, before step 210, that is, before the model to be trained is trained using the sample image, at least one of the encoder model and the decoder model may be compressed based on any one of a model pruning algorithm, a model distillation algorithm, and a model quantization algorithm.
[0112] In the present application, at least one of the encoder model and the decoder model is compressed. The advantage is that the model volume of the encoder model and the decoder model can be further reduced with little or no loss in model accuracy, thereby further accelerating the calculation speed and saving computer resources.
[0113] In such Fig. 9 In one embodiment of step 210, the model to be trained is trained by the sample image to obtain the information extraction model, which can be performed according to the steps shown in step 10:
[0114] See also Fig.10, shows a detailed flow chart of training the model to be trained by the sample image according to an embodiment of the present application. Specifically including steps 211 to 214:
[0115] Step 211: Encode the sample image through the encoder model to obtain sample encoding feature data.
[0116] Step 212: Decode the sample encoded feature data respectively through the at least two decoder models to obtain at least two groups of sample feature information, wherein each group of sample feature information includes feature information for each information unit in the sample image.
[0117] Step 213: trigger each of the at least two decoder models to learn sample feature information decoded by other decoder models except itself.
[0118] Step 214, determining a decoder model from the at least two decoder models as a target decoder model, and obtaining the information extraction model composed of the encoder model and the target decoder model.
[0119] In such Fig.10 In one embodiment of step 212, the sample encoding feature data is decoded by the at least two decoder models to obtain at least two sets of sample feature information. Fig.11 The steps shown are performed:
[0120] See also Fig.11 , shows a detailed flow chart of decoding the sample encoding feature data respectively by the at least two decoder models according to an embodiment of the present application. Specifically comprising steps 2121 to 2122:
[0121] Step 2121: Based on the sample coding feature data, the sample positioning information of each information unit in the sample image is determined by the at least two decoder models according to different arrangement directions of the information units in the sample image.
[0122] Step 2122: Based on the sample positioning information corresponding to the at least two decoder models, the sample encoding feature data are decoded respectively to obtain the at least two groups of sample feature information.
[0123] In order to enable those skilled in the art to better understand the above embodiments, the following will continue to take the formula recognition scene in the image as an example. Fig.12 Let's take a specific example to illustrate.
[0124] See also Fig.12, showing a framework diagram of a training information extraction model according to an embodiment of the present application.
[0125] like Fig.12 As shown, the model to be trained includes an encoder model 1201, a first decoder model 1202 and a second decoder model 1203.
[0126] First, the encoder model encodes the target image containing the formula "A+BC" to obtain encoded feature data, and then the first decoder model 1202 and the second decoder model 1203 respectively decode the encoded feature data based on the attention mechanism to obtain two sets of feature information.
[0127] In the process of decoding the encoded feature data, the first decoder model 1202 and the second decoder model 1203 can respectively determine the positioning information of each information unit in the target image according to the different arrangement directions of the information units (i.e., character units) in the target image, and then decode the encoded feature data based on the positioning information to obtain two sets of sample feature information.
[0128] For example, the first decoder model 1202 can determine the positioning information of the character "A", the character "+", the character "B", the character "-" and the character "C" according to the arrangement direction of "A" → "+" → "B" → "-" → "C" in the formula "A+BC", and the second decoder model 1203 can determine the positioning information of the character "C", the character "-", the character "B", the character "+" and the character "A" according to the arrangement direction of "C" → "-" → "B" → "+" → "A" in the formula "A+BC". After determining the positioning information of each character unit in the formula "A+BC", the first decoder model 1202 and the second decoder model 1203 can extract feature information of the character unit based on the positioning information of the character unit determined by themselves.
[0129] Furthermore, after obtaining two sets of feature information for the formula "A+BC", the first decoder model 1202 is triggered to learn the feature information decoded by the second decoder model 1203, and the second decoder model 1203 is also triggered to learn the feature information decoded by the first decoder model 1202, so as to optimize the model parameters in the first decoder model 1202 and the second decoder model 1203 respectively.
[0130] Finally, a decoder model is retained as a target decoder model between the first decoder model 1202 and the second decoder model 1203 (for example, the first decoder model 1202 is retained as the target decoder model), and the information extraction model composed of the encoder model and the target decoder model is obtained.
[0131] In the present application, by respectively performing decoding training on each information unit in the image according to different arrangement directions of the information units in the image, the at least two decoder models can simultaneously focus on and fully utilize the information of different angles of the information units in the image (such as historical and future information). Furthermore, through mutual learning, the decoder models can fully utilize the complementary information of different arrangement directions of the information units and explore long-distance dependency information, thereby making the decoding capability of the decoder model stronger, thereby helping to improve the accuracy of information extraction.
[0132] In this application, we continue to take the formula recognition scenario in the image as an example. Figure 2 After step 280, that is, after extracting the target feature information of the second information unit based on the target positioning information, the following steps may also be performed: Fig.13 Steps shown.
[0133] See also Fig.13 , shows a flowchart of a method after extracting target feature information of the second information unit based on the target positioning information according to an embodiment of the present application. Specifically comprising steps 291 to 292:
[0134] Step 291, obtaining target feature information corresponding to each character unit in the target image.
[0135] Step 292: based on the target feature information, edit one or more formulas in the target image into a formula editing area.
[0136] Specifically, in this application scenario, when editing a document, the user can capture the formula image that needs to be edited on the web page, and then, based on the information extraction scheme proposed in this application, extract the first feature information of at least one character unit in the formula image to obtain the target feature information, and then edit the formula in the formula image to the formula editing area based on the target feature information. It can be seen that the information extraction method proposed in this application can bring great convenience and excellent user experience to users in the formula editing process.
[0137] In the present application, at least two positioning reference information can be obtained by convolving the historical positioning information of the first information unit whose feature information has been extracted in the target image through at least two positioning convolution kernels of different scales, and then determining the target positioning information of the second information unit whose feature information has not been extracted in the target image based on the at least two positioning reference information, and finally extracting the target feature information of the second information unit based on the target positioning information. Since the historical positioning information is convolved with at least two positioning convolution kernels of different scales when determining the target positioning information of the second information unit, the positioning convolution kernels of different scales can be used to focus on information units of different sizes in the target image, so the obtained at least two positioning reference information also include the focus information for information units of different sizes, so that the target positioning information with higher accuracy for the second information unit can be determined based on the at least two positioning reference information, thereby improving the accuracy of the target feature information of the second information unit determined by the target positioning information.
[0138] The following describes an apparatus embodiment of the present application, which can be used to execute the information extraction method in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned embodiment of the information extraction method of the present application.
[0139] Fig.14 A block diagram of an information extraction device according to an embodiment of the present application is shown.
[0140] Reference Fig.14 As shown, according to an embodiment of the present application, an information extraction device 1400 includes: an acquisition unit 1401, a convolution unit 1402, a determination unit 1403, and an extraction unit 1404.
[0141] Among them, the acquisition unit 1401 is used to acquire the historical positioning information of the first information unit whose feature information has been extracted in the target image, and the target image includes at least one information unit; the convolution unit 1402 is used to perform convolution processing on the historical positioning information through at least two positioning convolution kernels of different scales, respectively, to obtain at least two positioning reference information; the determination unit 1403 is used to determine the target positioning information of the second information unit whose feature information has not been extracted in the target image based on the at least two positioning reference information; the extraction unit 1404 is used to extract the target feature information of the second information unit based on the target positioning information. .
[0142] In some embodiments of the present application, based on the aforementioned scheme, the acquisition unit 1401 is configured to: obtain the historical positioning information of each first information unit that has completed feature information extraction in the target image; calculate the historical positioning information of each first information unit that has completed feature information extraction to obtain the historical positioning information.
[0143] In some embodiments of the present application, based on the aforementioned scheme, the convolution unit 1402 is configured to: perform convolution processing on the historical positioning information through the first positioning convolution kernel in the information extraction model to obtain first positioning reference information; perform convolution processing on the historical positioning information through the second positioning convolution kernel in the information extraction model to obtain second positioning reference information, and the scale of the second positioning convolution kernel is larger than the scale of the first positioning convolution kernel.
[0144] In some embodiments of the present application, based on the aforementioned scheme, the determination unit 1403 is configured to: before determining the target positioning information of the second information unit in the target image for which feature information extraction has not been completed based on the at least two positioning reference information, determine the second information unit from the information units in the target image for which feature information extraction has not been completed according to a preset arrangement direction of the at least one information unit in the target image.
[0145] In some embodiments of the present application, based on the aforementioned scheme, the determination unit 1403 is also configured to: obtain the hidden state information of the information unit that is closest to the second information unit when the feature information is extracted, and obtain the encoded feature data for the target image; aggregate the at least two positioning reference information, the hidden state information, and the encoded feature data to obtain the target positioning information of the second information unit.
[0146] In some embodiments of the present application, based on the aforementioned scheme, the extraction unit 1404 is configured to: obtain the encoded feature data for the target image; based on the target positioning information, decode the encoded feature data through the target decoder model in the information extraction model to obtain the target feature information of the second information unit.
[0147] In some embodiments of the present application, based on the aforementioned scheme, the device also includes: a training unit, which is used to obtain a model to be trained before obtaining historical positioning information of a first information unit from which feature information extraction has been completed in a target image, the model to be trained comprising an encoder model and at least two decoder models, the encoder model being used to encode an image to obtain encoded feature data, and the decoder model being used to decode the encoded feature data to obtain feature information of each information unit in the image; obtaining a sample image, and training the model to be trained using the sample image to obtain an information extraction model.
[0148] In some embodiments of the present application, based on the aforementioned scheme, the encoder model includes a densely connected convolutional network model.
[0149] In some embodiments of the present application, based on the aforementioned scheme, the training unit is configured to: encode the sample image through the encoder model to obtain sample encoding feature data; decode the sample encoding feature data through the at least two decoder models respectively to obtain at least two groups of sample feature information, wherein each group of sample feature information includes feature information for each information unit in the sample image; trigger each of the at least two decoder models to learn the sample feature information decoded by other decoder models except itself; determine a decoder model as a target decoder model among the at least two decoder models to obtain the information extraction model composed of the encoder model and the target decoder model.
[0150] In some embodiments of the present application, based on the aforementioned scheme, the training unit is configured as: based on the sample coding feature data, determining the sample positioning information of each information unit in the sample image according to different arrangement directions of the information units in the sample image through the at least two decoder models; based on the sample positioning information corresponding to the at least two decoder models, respectively decode the sample coding feature data to obtain the at least two groups of sample feature information.
[0151] In some embodiments of the present application, based on the aforementioned scheme, the information unit includes a character unit, at least one information unit in the target image constitutes one or more formulas containing the character unit, and the at least two positioning reference information are used to respectively focus on the positioning information of character units of different sizes.
[0152] In some embodiments of the present application, based on the aforementioned scheme, the device also includes: an editing unit, which is used to obtain target feature information corresponding to each character unit in the target image after extracting the target feature information of the second information unit based on the target positioning information; based on the target feature information, edit one or more formulas in the target image to a formula editing area.
[0153] As another aspect, an embodiment of the present application also provides another information extraction device, including a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing the information extraction method described in the above embodiments.
[0154] Fig.15 A block diagram of an information extraction device according to an embodiment of the present application is shown. For example, the device 1500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0155] Reference Fig.15 , device 1500 may include one or more of the following components: a processing component 1502 , a memory 1504 , a power component 1506 , a multimedia component 1508 , an audio component 1510 , an input / output (I / O) interface 1512 , a sensor component 1514 , and a communication component 1516 .
[0156] The processing component 1502 generally controls the overall operation of the device 1500, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 1502 may include one or more processors 1518 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 1502 may include one or more modules to facilitate the interaction between the processing component 1502 and other components. For example, the processing component 1502 may include a multimedia module to facilitate the interaction between the multimedia component 1508 and the processing component 1502.
[0157] The memory 1504 is configured to store various types of data to support operations on the device 1500. Examples of such data include instructions for any application or method operating on the device 1500, contact data, phone book data, messages, pictures, videos, etc. The memory 1504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0158] The power supply component 1506 provides power to the various components of the device 1500. The power supply component 1506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 1500.
[0159] The multimedia component 1508 includes a screen that provides an output interface between the device 1500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1508 includes a front camera and / or a rear camera. When the device 1500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0160] The audio component 1510 is configured to output and / or input audio signals. For example, the audio component 1510 includes a microphone (MIC), and when the device 1500 is in an operation mode, such as a call mode, a recording mode, and a voice information processing mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 1504 or sent via the communication component 1516. In some embodiments, the audio component 1510 also includes a speaker for outputting audio signals.
[0161] I / O interface 1512 provides an interface between processing component 1502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0162] The sensor assembly 1514 includes one or more sensors for providing various aspects of the status assessment of the device 1500. For example, the sensor assembly 1514 can detect the open / closed state of the device 1500, the relative positioning of components, such as the display and keypad of the device 1500, the sensor assembly 1514 can also search for changes in the position of the search result display device 1500 or a component of the device 1500, the presence or absence of user contact with the device 1500, the orientation or acceleration / deceleration of the device 1500, and the temperature change of the device 1500. The sensor assembly 1514 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 1514 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1514 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0163] The communication component 1516 is configured to facilitate wired or wireless communication between the device 1500 and other devices. The device 1500 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1516 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency information processing (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0164] In an exemplary embodiment, the apparatus 1500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above methods.
[0165] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1504 including instructions, and the instructions can be executed by a processor 1518 of the device 1500 to perform the above-mentioned information extraction method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0166] As another aspect, the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the information extraction method described in the above embodiment.
[0167] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The computer-readable storage medium stores at least one program code, which is loaded and executed by the processor of the device to implement the operations performed by the information extraction method described in the above embodiment.
[0168] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0169] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.
[0170] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.
[0171] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An information extraction method, characterized in that: The method comprises: Acquire historical positioning information of a first information unit from a target image for which feature information extraction has been completed, wherein the target image includes at least one information unit; The historical positioning information is convolved using at least two positioning convolution kernels of different scales to obtain at least two positioning reference information; Based on the at least two positioning reference information, determine the target positioning information of the second information unit in the target image for which feature information extraction has not been completed; the target positioning information is obtained by aggregating the at least two positioning reference information, hidden state information, and encoded feature data for the target image, and the hidden state information is obtained when extracting feature information from an information unit that is closest to the second information unit; Based on the target positioning information, target feature information of the second information unit is extracted, wherein the target feature information is obtained by decoding the encoded feature data through a target decoder model in an information extraction model based on the target positioning information.
2. The method according to claim 1, characterized in that The step of obtaining historical positioning information of the first information unit in the target image from which feature information extraction has been completed includes: Acquire historical positioning information of each first information unit from which feature information extraction has been completed in the target image; The historical positioning information of each first information unit from which feature information extraction has been completed is calculated to obtain the historical positioning information.
3. The method according to claim 1, characterized in that The historical positioning information is convolved by at least two positioning convolution kernels of different scales to obtain at least two positioning reference information, including: Performing convolution processing on the historical positioning information through a first positioning convolution kernel in the information extraction model to obtain first positioning reference information; The historical positioning information is convolved by a second positioning convolution kernel in the information extraction model to obtain second positioning reference information, wherein the scale of the second positioning convolution kernel is larger than the scale of the first positioning convolution kernel.
4. The method according to claim 1, characterized in that: Before determining the target positioning information of the second information unit in the target image for which feature information extraction has not been completed based on the at least two positioning reference information, the method further includes: According to a preset arrangement direction of the at least one information unit in the target image, the second information unit is determined from information units in the target image for which feature information extraction has not been completed.
5. The method according to claim 1, characterized in that The determining, based on the at least two positioning reference information, the target positioning information of the second information unit in the target image for which feature information extraction has not been completed includes: Obtaining hidden state information of the information unit that is closest to the second information unit during feature information extraction, and obtaining coded feature data for the target image; The at least two positioning reference information, the hidden state information, and the encoding feature data are aggregated to obtain the target positioning information of the second information unit.
6. The method according to claim 1, characterized in that The extracting target feature information of the second information unit based on the target positioning information includes: Acquire encoding feature data for the target image; Based on the target positioning information, the encoded feature data is decoded by a target decoder model in an information extraction model to obtain target feature information of the second information unit.
7. The method according to claim 1, characterized in that Before acquiring the historical positioning information of the first information unit from which feature information extraction has been completed in the target image, the method further includes: Obtaining a model to be trained, wherein the model to be trained includes an encoder model and at least two decoder models, wherein the encoder model is used to encode an image to obtain encoded feature data, and the decoder model is used to decode the encoded feature data to obtain feature information of each information unit in the image; A sample image is obtained, and the model to be trained is trained using the sample image to obtain an information extraction model.
8. The method according to claim 7, characterized in that The encoder model includes a densely connected convolutional network model.
9. The method according to claim 7, characterized in that: The step of training the model to be trained by using the sample image to obtain an information extraction model includes: Encoding the sample image by using the encoder model to obtain sample encoding feature data; Decoding the sample encoded feature data respectively through the at least two decoder models to obtain at least two groups of sample feature information, wherein each group of sample feature information includes feature information for each information unit in the sample image; Triggering each of the at least two decoder models to learn sample feature information decoded by other decoder models except itself; A decoder model is determined as a target decoder model among the at least two decoder models, and the information extraction model consisting of the encoder model and the target decoder model is obtained.
10. The method according to claim 9, characterized in that The decoding of the sample encoding feature data by the at least two decoder models respectively to obtain at least two sets of sample feature information includes: Based on the sample coding feature data, determining sample positioning information of each information unit in the sample image according to different arrangement directions of the information units in the sample image by the at least two decoder models; Based on the sample positioning information corresponding to the at least two decoder models, the sample encoding feature data is decoded respectively to obtain the at least two groups of sample feature information.
11. The method according to any one of claims 1 to 10, characterized in that: The information unit includes a character unit, at least one information unit in the target image constitutes one or more formulas containing the character unit, and the at least two positioning reference information are used to respectively focus on the positioning information of character units of different sizes.
12. The method according to any one of claims 1 to 10, characterized in that: After extracting the target feature information of the second information unit based on the target positioning information, the method further includes: Acquire target feature information corresponding to each character unit in the target image; Based on the target feature information, one or more formulas in the target image are edited into a formula editing area.
13. An information extraction device, characterized in that: The device comprises: An acquisition unit, used to acquire historical positioning information of a first information unit from which feature information extraction has been completed in a target image, wherein the target image includes at least one information unit; A convolution unit is used to perform convolution processing on the historical positioning information through at least two positioning convolution kernels of different scales to obtain at least two positioning reference information; A determination unit is used to determine, based on the at least two positioning reference information, target positioning information of a second information unit in the target image for which feature information extraction has not been completed; the target positioning information is obtained by aggregating the at least two positioning reference information, hidden state information, and encoded feature data for the target image, and the hidden state information is obtained when extracting feature information from an information unit that is closest to the second information unit; The extraction unit is used to extract target feature information of the second information unit based on the target positioning information, wherein the target feature information is obtained by decoding the encoded feature data through a target decoder model in the information extraction model based on the target positioning information.
14. The device according to claim 13, characterized in that The acquisition unit is configured as follows: Acquire historical positioning information of each first information unit from which feature information extraction has been completed in the target image; The historical positioning information of each first information unit from which feature information extraction has been completed is calculated to obtain the historical positioning information.
15. The device according to claim 13, characterized in that The convolution unit is configured as follows: Performing convolution processing on the historical positioning information through a first positioning convolution kernel in the information extraction model to obtain first positioning reference information; The historical positioning information is convolved by a second positioning convolution kernel in the information extraction model to obtain second positioning reference information, wherein the scale of the second positioning convolution kernel is larger than the scale of the first positioning convolution kernel.
16. The device according to claim 13, characterized in that The determining unit is configured to: Before determining the target positioning information of the second information unit whose feature information extraction has not been completed in the target image based on the at least two positioning reference information, the second information unit is determined from the information units whose feature information extraction has not been completed in the target image according to a preset arrangement direction of the at least one information unit in the target image.
17. The device according to claim 13, characterized in that The determining unit is further configured to: Obtaining hidden state information of the information unit that is closest to the second information unit during feature information extraction, and obtaining coded feature data for the target image; The at least two positioning reference information, the hidden state information, and the encoding feature data are aggregated to obtain the target positioning information of the second information unit.
18. The device according to claim 13, characterized in that The extraction unit is configured to: Acquire encoding feature data for the target image; Based on the target positioning information, the encoded feature data is decoded by a target decoder model in an information extraction model to obtain target feature information of the second information unit.
19. The device according to claim 13, characterized in that The device also includes a training unit, which is used to: Before obtaining historical positioning information of a first information unit from which feature information extraction has been completed in a target image, obtaining a model to be trained, wherein the model to be trained includes an encoder model and at least two decoder models, the encoder model is used to encode the image to obtain encoded feature data, and the decoder model is used to decode the encoded feature data to obtain feature information of each information unit in the image; A sample image is obtained, and the model to be trained is trained using the sample image to obtain an information extraction model.
20. The device according to claim 19, characterized in that The encoder model includes a densely connected convolutional network model.
21. The device according to claim 19, characterized in that The training unit is configured as follows: Encoding the sample image by using the encoder model to obtain sample encoding feature data; Decoding the sample encoded feature data respectively through the at least two decoder models to obtain at least two groups of sample feature information, wherein each group of sample feature information includes feature information for each information unit in the sample image; Triggering each of the at least two decoder models to learn sample feature information decoded by other decoder models except itself; A decoder model is determined as a target decoder model among the at least two decoder models, and the information extraction model consisting of the encoder model and the target decoder model is obtained.
22. The device according to claim 21, characterized in that The training unit is configured as follows: Based on the sample coding feature data, determining sample positioning information of each information unit in the sample image according to different arrangement directions of the information units in the sample image by the at least two decoder models; Based on the sample positioning information corresponding to the at least two decoder models, the sample encoding feature data is decoded respectively to obtain the at least two groups of sample feature information.
23. The device according to any one of claims 13 to 22, characterized in that The information unit includes a character unit, at least one information unit in the target image constitutes one or more formulas containing the character unit, and the at least two positioning reference information are used to respectively focus on the positioning information of character units of different sizes.
24. The device according to any one of claims 13 to 22, characterized in that The device further comprises an editing unit: configured to obtain target feature information corresponding to each character unit in the target image after extracting target feature information of the second information unit based on the target positioning information; Based on the target feature information, one or more formulas in the target image are edited into a formula editing area.
25. An information extraction device, characterized in that: The invention comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs include instructions for performing the information extraction method according to any one of claims 1 to 12.
26. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by a processor to implement the operations performed by the information extraction method according to any one of claims 1 to 12.
27. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are suitable for being read and executed by a processor, so that a computer device having the processor executes the information extraction method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Positioning method and device, electronic equipment and computer readable storage medium
CN111417066A
Neural network model for image segmentation and image segmentation method therefor
WO2021128896A1