Information extraction method and device and electronic equipment

By splitting the card and document recognition task into multiple subtasks for parallel processing and using an attention mask matrix to isolate the subtasks, the efficiency of card and document information extraction is improved, the problem of long processing time for complex card and document recognition is solved, and rapid information extraction is achieved.

CN121640494APending Publication Date: 2026-03-10VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies have low information extraction efficiency in card and document recognition scenarios, especially for complex cards and documents such as driver's licenses and household registration books, which takes a long time and is difficult to meet user needs.

Method used

By obtaining the configuration file corresponding to the image category, the information extraction task is split into multiple subtasks, and the field information of each information field is processed in parallel. The attention mask matrix is ​​used to isolate the subtasks, thereby improving the information extraction efficiency.

Benefits of technology

The information retrieval time has been optimized, reducing it from 3-4 seconds for complex cards to within 1 second, and the total time has been reduced from 5-6 seconds to 3 seconds, meeting user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640494A_ABST
    Figure CN121640494A_ABST
Patent Text Reader

Abstract

The invention discloses an information extraction method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The information extraction method comprises the steps that a configuration file corresponding to the image category of a first image is acquired, the first image comprises a plurality of information fields, the configuration file comprises a task cue word, and the task cue word is used for prompting extraction of field information of all the information fields in the first image; according to the information fields, processing the task cue word to obtain a sub-task cue word corresponding to each information field, the sub-task cue word being used for prompting to extract field information of one information field; and outputting field information of each information field according to each sub-task cue word.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to an information extraction method and device and electronic equipment. BACKGROUND

[0002] Large models are increasingly used on electronic devices, and many scenarios such as text summarization and input polishing will use large models. As user scenario requirements continue to increase, some scenarios have high requirements for the performance of large models.

[0003] For example, in the card certificate recognition scenario, a multi-modal large model is currently mainly used on the terminal side for recognition. First, a complex card certificate image is processed, then the multi-modal large model is input, and finally the recognition result is output. The entire process needs to return to the user within 3s, and the experience is good. The current complex card certificates cannot meet this requirement. For example, a driver's license and a household booklet, which have a large number of field types to be extracted, such as more than 8 fields of license plate number, vehicle type, owner, and vehicle identification code on a driver's license. The information of all fields adds up to several tens to hundreds of words. The household booklet type can even reach 150-200 words. According to the current information extraction scheme, the time consumption of only outputting words reaches 3s-4s, and the total time consumption of the whole process reaches nearly 6s. The information extraction efficiency is low, and it is difficult to meet the user's needs. SUMMARY

[0004] The purpose of the embodiments of the application is to provide an information extraction method, device and electronic equipment, which can improve the information extraction efficiency.

[0005] In a first aspect, the embodiments of the application provide an information extraction method, comprising: obtaining a configuration file corresponding to an image category of a first image, the first image comprising a plurality of information fields, the configuration file comprising a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image; processing the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word being used to prompt extraction of field information of one information field; outputting the field information of each information field according to the sub-task prompt words.

[0006] In a second aspect, the embodiments of the application provide an information extraction device, comprising: an obtaining module, configured to obtain a configuration file corresponding to an image category of a first image, the first image comprising a plurality of information fields, the configuration file comprising a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image; The processing module is configured to process the task prompt word according to the information field, to obtain a sub-task prompt word corresponding to each information field, and the sub-task prompt word is used to prompt extraction of field information of one information field. The output module is configured to output the field information of each information field according to the sub-task prompt words.

[0007] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. When the programs or instructions are executed by the processor, the steps of the method according to the first aspect are implemented.

[0008] In a fourth aspect, a readable storage medium is provided. The readable storage medium stores programs or instructions. When the programs or instructions are executed by a processor, the steps of the method according to the first aspect are implemented.

[0009] In a fifth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is configured to execute programs or instructions, and implement the steps of the method according to the first aspect.

[0010] In a sixth aspect, a computer program product is provided. The program product is stored in a storage medium. When the program product is executed by at least one processor, the steps of the method according to the first aspect are implemented.

[0011] According to the image category of the image, the pre-set configuration file can be obtained. Based on the configuration file, the task prompt word corresponding to the image can be determined. Based on the task prompt word, the sub-task prompt words corresponding to the information fields can be obtained. The task of extracting multiple information fields is split into multiple sub-tasks. Based on the sub-task prompt words of the multiple sub-tasks, the field information of each information field can be predicted in parallel. The information extraction efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 A flowchart of an information extraction method provided by an embodiment of the present application is provided. Figure 2 A flowchart of another information extraction method provided by an embodiment of the present application is provided. Figure 3 A flowchart of another information extraction method provided by an embodiment of the present application is provided. Figure 4 A flowchart of another information extraction method provided by an embodiment of the present application is provided. Figure 5 A partial schematic diagram of a word output stage provided by an embodiment of the present application is provided. Figure 6A schematic diagram of three output token sequences provided by an embodiment of the present application is shown in the following table. Figure 7 A schematic diagram of a first attention mask matrix provided by an embodiment of the present application is shown in the following table. Figure 6 Figure 8 A schematic diagram of a second attention mask matrix provided by an embodiment of the present application is shown in the following table. Figure 9 A schematic diagram of a second attention mask matrix provided by an embodiment of the present application is shown in the following table. Figure 10 A schematic diagram of a second attention mask matrix provided by an embodiment of the present application is shown in the following table. Figure 11 A schematic diagram of a second attention mask matrix provided by an embodiment of the present application is shown in the following table. Figure 12 A schematic diagram of a second attention mask matrix provided by an embodiment of the present application is shown in the following table. DETAILED DESCRIPTION

[0013] The technical solutions in the embodiments of the present application will be described clearly below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0014] The terms "first", "second", and the like in the specification of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a class, not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / ", generally represents a "or" relationship between the front and rear associated objects.

[0015] When a certain image contains multiple information fields, in order to extract the field information of each information field, the related technology mainly uses a large language model to perform serial processing on the image, resulting in low word output efficiency, long time consumption of the entire process, and low information extraction efficiency.

[0016] Therefore, the embodiments of the present application provide an information extraction method, device and electronic equipment, which can improve the information extraction efficiency.

[0017] The information extraction method, device and electronic equipment provided by the embodiments of the present application will be described in detail below with reference to the drawings and through specific embodiments and application scenarios.​

[0018] Figure 1 A flowchart of an information extraction method provided by an embodiment of the present application. The information extraction method can be applied to electronic devices such as mobile phones, tablets, and laptops. As shown in Figure 1 The information extraction method can include the following S10-S130: S110, obtaining a configuration file corresponding to the image category of the first image, the first image including multiple information fields, the configuration file including a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image.

[0019] The first image described above may, for example, be a card image, a bill image, a certificate image, a form image, or the like containing multiple information fields. Taking a driver's license image as an example, the driver's license image may, for example, include a license number, a name, an address, a first license date, and the like.

[0020] In the present embodiment, after obtaining the first image, the image category of the first image can be determined. The image category may, for example, include categories such as driver's license, ID card, xx certificate, and xx form. The first image may, for example, be input into a category recognition model to output the image category of the first image. The category recognition model may, for example, include a deep learning model. The features of the first image may, for example, also be extracted, and then classified by a machine learning algorithm to obtain the category of the first image. Of course, other ways of recognizing the image category of the first image can also be used.

[0021] Different image categories can correspond to different configuration files, which can include a task prompt word for prompting extraction of field information of all information fields in the first image. That is, different image categories can correspond to different task prompt words. By determining the image category of the first image, the corresponding task prompt word can be accurately obtained, providing a more accurate basis for subsequent extraction of field information of each information field.

[0022] The configuration file can be stored locally, and the corresponding configuration file can be obtained from the local based on the image category of the first image.

[0023] The task prompt word is used to prompt extraction of field information of all information fields in the first image. The task prompt word may, for example, include a basic prompt word and a field prompt word. The basic prompt word may, for example, include "extract the field information of A and convert it to json format, only return:". Here, A represents the image category of the first image, for example, A can be a driver's license, an ID card, a form, a bill, and the like.

[0024] The field prompt words are used to prompt each information field in the first image. For example, for a driver's license, the field prompt words can include "[license number] field information:", "[name] field information:", "[address] field information:", "[date of birth] field information:", and "[date of first license issue] field information:".

[0025] In some embodiments, the complete form of the task prompt word can be as follows: "base_prompt": { "category 1": "[special_token_start] Extract the field information of A and convert it into json format, only need to return:"} "card_prompt": { "category 1": [ "[license number] field information: [special_token_end]:", "[name] field information: [special_token_end]:", "[address] field information: [special_token_end]:", "[date of birth] field information: [special_token_end]:", "[date of first license issue] field information: [special_token_end]:"} Wherein, the base_prompt represents the basic prompt word, the card_prompt represents the field prompt word, and [special_token_start] and [special_token_end] are used to represent the start and end of the task.

[0026] S120, processing the task prompt word according to the information field, to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word is used to prompt the extraction of the field information of one information field.

[0027] The processing here mainly splits the task prompt word, splits a task of extracting field information containing multiple information fields into multiple sub-tasks, one sub-task corresponding to one information field, provides a basis for subsequent parallel processing of multiple sub-tasks, and thus improves the information extraction efficiency. The specific splitting process can be referred to in the following embodiments.

[0028] Each sub-task prompt word can include a basic prompt word and a corresponding field prompt word, so as to ensure the integrity of the sub-task prompt word and play an accurate prompting role.

[0029] S130, outputting the field information of each information field according to the sub-task prompt words.

[0030] Exemplarily, after obtaining the sub-task prompt words corresponding to each information field, the sub-task prompt words of each information field can be spliced to obtain a spliced prompt word, which provides a basis for subsequent field information prediction. In order to realize parallel processing of multiple sub-tasks, a separator can be added between the sub-task prompt words when splicing the sub-task prompt words of each information field. The form of the separator is not limited in the embodiment, for example, "+" can be used to represent.

[0031] In the embodiment, in order to facilitate subsequent information extraction, the spliced prompt word can be converted into a token sequence. Exemplarily, the spliced prompt word can be subjected to word segmentation processing, for example, the spliced prompt word can be input into a word segmentor, and the word segmentor can perform word segmentation processing on the spliced prompt word. Then, in combination with a token mapping relationship table, the token corresponding to each word segmentation result can be obtained, and then the token sequence corresponding to the spliced prompt word can be obtained by sequentially combining the word segmentation results.

[0032] Exemplarily, the token sequence corresponding to the spliced prompt word can be input into an information prediction model, and each information field and the field information of each information field can be output by the information prediction model.

[0033] The embodiment of the application can obtain a pre-set configuration file according to the image category of the first image, determine the task prompt word corresponding to the first image based on the configuration file, and obtain the sub-task prompt word corresponding to each information field based on the task prompt word, thereby realizing the splitting of one task of extracting multiple information fields into multiple sub-tasks. The field information of each information field can be predicted in parallel based on the sub-task prompt words of the multiple sub-tasks, thereby improving the information extraction efficiency.

[0034] In order to more accurately extract the field information of each information field, Figure 2 Exemplarily, a flowchart of an information extraction method is provided, Figure 2 which is different from Figure 1 in that Figure 1 S130 in the embodiment can be refined as Figure 2 S210-S220 in the embodiment.

[0035] S210, splicing each sub-task prompt word to obtain a spliced prompt word.

[0036] The specific splicing process can be referred to the above-mentioned embodiments, which will not be described here again.

[0037] S220, inputting the token sequence corresponding to the spliced prompt word and the attention mask matrix into an information prediction model to output each information field and the field information of each information field.

[0038] The attention mask matrix is used to isolate the subtask prompt words corresponding to each information field of the input and isolate the field information of each information field of the output.

[0039] The attention mask matrix can be a matrix containing 0 and inf, where inf represents infinity, and 0 represents allowing the calculation of attention distribution, and inf represents prohibiting the calculation of attention distribution. In this embodiment, the attention mask matrix can be multiple, for example, it can include a first attention mask matrix for isolating the subtask prompt words corresponding to each information field of the input, and a second attention mask matrix for isolating the field information of each information field of the output. The determination process of the first attention mask matrix and the second attention mask matrix can refer to the following embodiments.

[0040] Since each information field is independent of each other, the parallel execution of multiple subtasks can be realized through the above attention mask matrix, which can avoid the interference of other information fields while improving the information extraction efficiency, thereby improving the accuracy of the information extraction result.

[0041] Exemplarily, the token sequence corresponding to the splicing prompt word and the attention mask matrix are input into the information prediction model, and the information prediction model outputs each information field and the field information of each information field. The information prediction model here can be a model using attention mechanism, and the specific structure of the model is not limited in this embodiment.

[0042] Moreover, the attention mask matrix can isolate the subtask prompt words corresponding to each information field of the input and isolate the field information of each information field of the output, that is, only calculate the attention distribution with the same information field in the information extraction process, avoiding the interference of other information fields and improving the accuracy of the information extraction result.

[0043] Taking the task prompt words including the basic prompt word and the field prompt word as an example, Figure 3 Exemplarily, a flowchart of an information extraction method is provided, Figure 3 Different from Figure 1 The difference is that Figure 1 S120 in can be refined to Figure 3 S310-S320 in.

[0044] S310, respectively extracting the basic prompt word and the field prompt word of each information field from the task prompt word.

[0045] Taking a driving license image as an example, exemplarily, a basic prompt word "extract driving license information and convert it into a json format, only return:" can be extracted from a task prompt word corresponding to the driving license image, and field prompt words of respective information fields are respectively "[license number] field information:", "[name] field information:", "[address] field information:", "[date of birth] field information:", and "[date of first license issuance] field information:".

[0046] S320, generating a sub-task prompt word corresponding to each information field according to the basic prompt word and the field prompt word of each information field.

[0047] Exemplarily, the basic prompt word can be combined with the field prompt word of each information field to obtain a sub-task prompt word corresponding to each information field.

[0048] For example, for the information field name, a sub-task prompt word "[special_token_start] extract driving license information and convert it into a json format, only return: [name] field information [special_token_end]" can be obtained, and similarly, for the information field license number, a sub-task prompt word "[special_token_start] extract driving license information and convert it into a json format, only return: [license number] field information [special_token_end]" can be obtained.

[0049] The embodiment can extract a basic prompt word and field prompt words of respective information fields from a task prompt word based on a configuration file, and obtain a sub-task prompt word corresponding to each information field based on the basic prompt word and the field prompt word of each information field, thereby providing a basis for task splitting, so that subsequent parallel processing of multiple sub-tasks can be performed, and information extraction efficiency is improved.

[0050] In order to parallel process each sub-task, in some embodiments, after S320, the information extraction method can further include the following steps: The basic prompt word and the field prompt word of each information field are spliced by a separator to obtain a spliced prompt word.

[0051] The separator here is used to splice sub-task prompt words of respective information fields, and simultaneously represents that information extraction tasks of respective information fields are independent of each other.

[0052] In some embodiments, the sub-task prompt words corresponding to respective information fields can be directly spliced by the separator to obtain the following spliced prompt word: "[special_token_start] extract the driving license information and convert it into json format, only need to return: [license number] field information [special_token_end] + [special_token_start] extract the driving license information and convert it into json format, only need to return: [name] field information [special_token_end] + [special_token_start] extract the driving license information and convert it into json format, only need to return: [address] field information [special_token_end] + [special_token_start] extract the driving license information and convert it into json format, only need to return: [date of birth] field information [special_token_end] + [special_token_start] extract the driving license information and convert it into json format, only need to return: [initial license date] field information [special_token_end]".

[0053] Since the subtask prompt words corresponding to each information field all contain the basic prompt word, in order to reduce the amount of calculation, the splicing prompt word can be optimized exemplarily, and the following splicing prompt word is obtained: "[special_token_start] extract the driving license information and convert it into json format, only need to return: [license number] field information [special_token_end] + [name] field information [special_token_end] + [address] field information [special_token_end] + [date of birth] field information [special_token_end] + [initial license date] field information [special_token_end]".

[0054] That is, the optimized splicing prompt word contains only one basic prompt word, and the basic prompt word and the field prompt words corresponding to each information field are separated by a separator, that is, the optimized splicing prompt word is actually composed of five independent subtask prompt words.

[0055] The embodiment can splice the subtask prompt words of each subtask through the separator, indicating that each subtask is independent of each other, and the information prediction model can execute each subtask in parallel after inputting the splicing prompt word into the information prediction model. In this way, the information extraction efficiency can be improved, and the amount of calculation can be reduced, and the computing resources can be saved.

[0056] Taking the attention mask matrix including the first attention mask matrix and the second attention mask matrix as an example, the first attention mask matrix is used for isolating the sub-task prompt words corresponding to each information field of the input, and the second attention mask matrix is used for isolating the field information of each information field of the output.

[0057] As shown in Figure 4 , Figure 4 With Figure 2 the difference is that, Figure 2 S220 in can be refined to S410-S430 in Figure 4 .

[0058] S410, input the token sequence corresponding to the splicing prompt word and the first attention mask matrix into the first information prediction module of the information prediction model, output the first token sequence, and the first token sequence includes one output token corresponding to each information field.

[0059] The first information prediction module is used for predicting the first output token of each information field, and the first output token here is the first token after the end of each field prompt word. Assuming that the license number of the driver's license of user Zhang XX is 1xxxx, the first token can be 1 for the information field "license number", and the first token can be "Zhang" for the information field "name". The first token sequence is a sequence generated based on the first output token of each information field. For example, the arrangement order of the information fields in the splicing prompt word is license number, name, address, date of birth, and date of first license, and the first output token of each information field in the first token sequence is also arranged in the same order.

[0060] The first attention mask matrix is used to ensure that the sub-task prompt words of each information field can only see the tokens inside their own field and cannot see the sub-task prompt words of other fields. At the same time, the self-recurrence characteristic is maintained inside each information field, that is, only the current token and the token before it can be seen, and the token after it cannot be seen. Exemplarily, the first attention mask matrix can be generated based on the token corresponding to the splicing prompt word, and the specific generation process can be referred to in the following embodiments.

[0061] In this embodiment, exemplarily, the token sequence corresponding to the splicing prompt word and the first attention mask matrix can be input into the first information prediction module, and the first output token of each information field is predicted in parallel through the first information prediction module.

[0062] Exemplarily, after the first word reasoning ends, one token of each information field of the reasoning can be spliced to obtain a first token sequence, and a length of the first token sequence is same as a number of the information fields.

[0063] S420, input the second token sequence, the token sequence corresponding to the concatenation prompt word, and the second attention mask matrix into a second information prediction module of the information prediction model, output a third token sequence, and each information field corresponding token in the third token sequence is a next token of each information field corresponding token in the second token sequence. The second token sequence includes the first token sequence.

[0064] The second token sequence is an output token corresponding to each information field that has been predicted, and the second token sequence includes the first token sequence. In the first iteration, the second token sequence is the first token sequence. That is, after obtaining the first token sequence, the first token sequence, the token sequence corresponding to the concatenation prompt word, and the second attention mask matrix are input into the second information prediction module of the information prediction model, so that a second output token of each information field can be predicted. At this time, the second token sequence includes the first output token and the second output token of each information field.

[0065] In the second iteration, the second token sequence, the token sequence corresponding to the concatenation prompt word, and the second attention mask matrix can be continuously input into the second information prediction module of the information prediction model, so that a third output token of each information field can be predicted. At this time, the second token sequence can include the first output token, the second output token, and the third output token of each information field. That is, the second token sequence is constantly updated based on the number of iterations. By analogy, all output tokens of each information field can be predicted.

[0066] Exemplarily, as Figure 5As shown, after obtaining the first token sequence 501, the token sequence 502 corresponding to the spliced prompt and the first token sequence 501 can be input into the second information prediction module 503, the second token of each information field is predicted through the second information prediction module 503, and the second token of each information field is spliced in order to obtain the second token sequence 504. Each token in each token sequence is an independent result. The 0, 1, 2, 3, and 4 in the first token sequence 501 and the second token sequence 504 correspond to different information fields, respectively. For example, when the first token sequence 501 includes the first output token of each information field, the second token sequence 504 includes the second output token of each information field.

[0067] After obtaining the second token, the token sequence corresponding to the spliced prompt, the predicted first token and the second token of each information field can continue to be input into the second information prediction module, the third token of each information field is predicted through the second information prediction module, and the same is true for the subsequent tokens. Among them, the token corresponding to each information field is input in the form of a sequence, each sequence contains the token at the same position of each information field, for example, the first token sequence contains the first token of each information field, the second token sequence contains the second token of each information field, the third token sequence contains the third token of each information field, and so on.

[0068] The second attention mask matrix is used to ensure that the generation process of each information field is independent when generating the subsequent token of each information field, that is, the generation of each information field can only see the sub-task prompt of its own field and the token of the field that has been generated, but cannot see the sub-task prompt of other fields and the token of other fields generated. At the same time, within each field, the generation process is autoregressive, that is, when generating the current token, only the prompt of the field and the token of the field generated before can be seen. Exemplarily, the second attention mask matrix can be generated based on the number of information fields, and the specific generation process can be referred to in the following embodiments.

[0069] S430, according to all token sequences predicted by the information prediction model, output each information field and field information of each information field.

[0070] Each token sequence corresponds to an output token of each information field, and based on all token sequences, the field information of each information field can be obtained.

[0071] The embodiment first combines the splicing prompt word and the first attention mask matrix to infer the first word of the field information, and on this basis, combines the splicing prompt word, the inferred first word, and the second attention mask matrix to infer the subsequent token of each information field, and obtains the field information of each information field based on the inferred token sequence, realizes parallel processing of each subtask, improves the information extraction efficiency, and at the same time, combines the first attention mask matrix and the second attention mask matrix to shield the interference of other information fields, and improves the accuracy of the information extraction result.

[0072] In some embodiments, the above S430 can include the following steps: Splitting each token sequence to obtain the token corresponding to each information field in each token sequence; for the same information field, splicing the tokens obtained from the token sequences to obtain the information field and the field information of the information field.

[0073] Taking three token sequences as an example, as shown in Figure 6 , it is assumed that the first token sequence is “1”, “Zhang”, “Beijing”, “1”, “2”, the second token sequence is “3”, “San”, “A”, “9”, “0”, and the third token sequence is “0”, “ ”, “City”, “9”, “1”. In actual application, there will be multiple token sequences. Among them, s0-s4 represent different information fields.

[0074] Exemplarily, each token sequence can be split to obtain the token corresponding to each information field in each token sequence, and for the same information field, the tokens obtained from the token sequences can be spliced, and the field information of each information field can be obtained in combination with the token mapping relationship table.

[0075] Exemplarily, as shown in Figure 7 , the field information of the certificate number s0 can be obtained as “130…”, the field information of the name s1 can be obtained as “Zhang San”, the field information of the address s2 can be obtained as “Beijing A City…”, the field information of the birth date s3 can be obtained as “199…”, and the field information of the initial certificate date s4 can be obtained as “201…”.

[0076] The final output result is: [the field information of the certificate number is: 130…], [the field information of the name is: Zhang San], [the field information of the address is: Beijing A City…], [the field information of the birth date is: 199…], and [the field information of the initial certificate date is: 201…].

[0077] After obtaining all token sequences, the token sequences are split, and the tokens at the same position of each token sequence are spliced, so that the field information of each information field is obtained, the parallel output of each information field is realized, and the information extraction efficiency is improved.

[0078] The generation process of the first attention mask matrix will be described below.

[0079] In some embodiments, after S110, the information extraction method can further include the following steps: According to the configuration file, the number of information fields contained in the first image is determined, and according to the field prompt words in the task prompt words, the tokens corresponding to each field prompt word are determined; According to the number of information fields and the tokens corresponding to each field prompt word, the first attention mask matrix is generated, and the number of rows and the number of columns of the first attention mask matrix are the same as the total number of tokens corresponding to each field prompt word; According to the position of each element in the first attention mask matrix and the information field to which the element corresponds, the element value of each element is determined; Among them, the element values of the lower triangular elements of the token matrix corresponding to the same information field in the first attention mask matrix are all first reference values, and the element values of the remaining elements are all second reference values, the first reference value indicates that the attention distribution is allowed to be calculated, and the second reference value indicates that the attention distribution is prohibited to be calculated.

[0080] Since each information field is independent of each other, in order to avoid interference between information fields, the first attention mask matrix can be generated based on the number of information fields and the tokens corresponding to each information field.

[0081] Taking an example of including two information fields, assuming that information field 1 corresponds to 6 tokens and information field 2 corresponds to 4 tokens, the first attention mask matrix is a 10*10 matrix.

[0082] The element value of each element in the first attention mask matrix can be determined according to the position of each element and the information field to which the element corresponds, so that the sub-task prompt words of each information field can only see the tokens inside their own field and cannot see the sub-task prompt words of other fields. At the same time, the self-recurrence characteristic is maintained inside each information field.

[0083] Taking the 6 tokens of information field 1 as t0-t5 and the 3 tokens of information field 2 as t6-t9 as an example, the first attention mask matrix can be as follows Figure 8The first attention mask matrix is shown in the table. In the first attention mask matrix, the element values of the lower triangular elements of the token matrix corresponding to the same information field are the first reference value, and the element values of the remaining elements are the second reference value.

[0084] Since information field 2 is independent of information field 1, information field 2 does not need to calculate the attention distribution with information field 1. Since t0-t5 belong to information field 1, the mask of information field 1 is normally calculated, and t5-t9 belong to information field 2, which does not need to calculate the mask with information field 1, so the values of t0-t5 positions are set to inf from the 7th row, and t6-t9 are calculated according to the normal mask. Thus, the second attention mask matrix can be obtained as shown in the table. Figure 8

[0085] The first attention mask matrix is generated based on the number of information fields and the tokens contained in each information field, and the relationship between the information fields is combined to modify the traditional attention mask matrix (all lower triangular elements are 0). Based on the modified attention mask matrix, only the attention distribution within the same information field needs to be calculated, and the attention distribution between other information fields does not need to be calculated, which simplifies the calculation amount, realizes the parallel processing of the reading sub-tasks, and improves the information extraction efficiency.

[0086] The generation process of the second attention mask matrix will be described below.

[0087] In some embodiments, after S110, the information extraction method can further include the following steps: determining the number of information fields contained in the first image according to the configuration file; generating a second attention mask matrix according to the number of information fields, the number of rows and the number of columns of the second attention mask matrix being the same as the total number of information fields, each row of the second attention mask matrix corresponding to an information field, and each column corresponding to an information field; determining the element value of each element in the second attention mask matrix according to the information fields corresponding to the row and column of each element; In the second attention mask matrix, the element value of the diagonal element is a third reference value, and the element value of the remaining element is a fourth reference value, the third reference value indicating that the attention distribution is allowed to be calculated, and the fourth reference value indicating that the attention distribution is prohibited to be calculated.

[0088] ​Since there is no context relationship between the tokens in each information field in the word output stage, that is, these tokens belong to different information fields respectively, the attention distribution between these tokens cannot be calculated. In order to obtain correct word output results and improve the accuracy of information extraction results, the second attention mask matrix is set in this embodiment.

[0089] In this embodiment, the second attention mask matrix is a mask matrix between the output tokens of each information field, and therefore can be generated based on the number of information fields, that is, the number of rows and columns of the second attention mask matrix is the same as the total number of information fields. Taking an example of including 5 information fields, the second attention mask matrix is exemplarily a 5*5 matrix, and each row corresponds to an information field.

[0090] Exemplarily, as shown in Figure 9 Since there is no context relationship between the 5 output tokens m0-m4, m0 does not need to calculate the attention distribution with m1-m4, similarly, m1 does not need to calculate the attention distribution with m0, m2-m4, m2 does not need to calculate the attention distribution with m0-m1, m3-m4, m3 does not need to calculate the attention distribution with m0-m2, m4, and m4 does not need to calculate the attention distribution with m0-m3. That is, the element value of the element on the diagonal line of the second attention mask matrix is 0, and the element value of the remaining elements is inf. In some embodiments, the third reference value is 0, and the fourth reference value is inf.

[0091] Therefore, the effect of shielding all other output tokens by each output token is achieved, the isolation of the output results of each information field is achieved, and the accuracy of the output results is improved. The above m0-m4 respectively represent the output token of each information field at a certain time.

[0092] The parallel inference scheme based on the large model provided in this embodiment can greatly optimize the word output performance of the first image, especially for complex cards such as household registration books and driving licenses, the word output time is optimized from 3-4s to within 1s, and the final total extraction time is optimized from 5-6s to within 3s, which can meet the needs of users.

[0093] It should be noted that the information extraction method provided in the embodiments of the present application can be executed by an information extraction device or a processing module in the information extraction device for executing the information extraction method. In the embodiments of the present application, the information extraction device executes the information extraction method as an example to illustrate the information extraction device provided in the embodiments of the present application.

[0094] Figure 10 The structure diagram of an information extraction device provided in the embodiments of the present application.

[0095] As Figure 10 shown, the information extraction apparatus 1000 can include: an acquisition module 1001 configured to acquire a configuration file corresponding to an image category of a first image, the first image including a plurality of information fields, the configuration file including a task prompt word, the task prompt word being used to prompt extraction of field information of all the information fields in the first image; a processing module 1002 configured to process the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word being used to prompt extraction of field information of one information field; an output module 1003 configured to output the field information of each information field according to the sub-task prompt words.

[0096] The embodiments of the present application can acquire a pre-set configuration file according to the image category of the first image, determine the task prompt word corresponding to the first image based on the configuration file, and obtain the sub-task prompt word corresponding to each information field based on the task prompt word, thereby achieving splitting of one task of extracting a plurality of information fields into a plurality of sub-tasks, and predicting the field information of each information field in parallel based on the sub-task prompt words of the plurality of sub-tasks and the attention mask matrix, thereby improving the information extraction efficiency.

[0097] In some possible implementations of the embodiments of the present application, the output module 1003 is specifically configured to: splice the sub-task prompt words to obtain a spliced prompt word; input the token sequence corresponding to the spliced prompt word and the attention mask matrix into an information prediction model, and output each information field and the field information of each information field; wherein the attention mask matrix is used to isolate the input sub-task prompt words corresponding to each information field and isolate the output field information of each information field.

[0098] In some possible implementations of the embodiments of the present application, the task prompt word includes a basic prompt word and a field prompt word of each information field; The processing module 1002 is specifically configured to: extract the basic prompt word and the field prompt word of each information field from the task prompt word; generate the sub-task prompt word corresponding to each information field according to the basic prompt word and the field prompt word of each information field.

[0099] In some possible implementations of the embodiments of the present application, the attention mask matrix includes a first attention mask matrix and a second attention mask matrix, the first attention mask matrix being used to isolate the input sub-task prompt words corresponding to each information field, and the second attention mask matrix being used to isolate the output field information of each information field; The output module 1003 is specifically configured to: input the token sequence corresponding to the splicing prompt word and the first attention mask matrix into a first information prediction module of the information prediction model, output a first token sequence, and the first token sequence includes one output token corresponding to each information field; input the second token sequence, the token sequence corresponding to the splicing prompt word, and the second attention mask matrix into a second information prediction module of the information prediction model, and output a third token sequence; the token corresponding to each information field in the third token sequence is the next token of the token corresponding to each information field in the second token sequence, the second token sequence is the output token corresponding to each information field that has been predicted, and the second token sequence includes the first token sequence; output each information field and field information of each information field according to all token sequences predicted by the information prediction model.

[0100] In some possible implementations of the embodiments of the present application, the attention mask matrix includes a first attention mask matrix, and the first attention mask matrix is used to isolate the input subtask prompt words corresponding to each information field; The processing module 1002 is further configured to, after the acquisition module acquires the configuration file corresponding to the image category of the first image, determine the number of information fields contained in the first image according to the configuration file, and determine the token corresponding to each field prompt word according to the field prompt word in the task prompt word; generate a first attention mask matrix according to the number of information fields and the token corresponding to each field prompt word, and the number of rows and the number of columns of the first attention mask matrix are the same as the total number of tokens corresponding to the field prompt words; determine the element value of each element according to the position of each element in the first attention mask matrix and the information field to which the token corresponding to the element belongs; In the first attention mask matrix, the element values of the lower triangular elements of the token matrix corresponding to the same information field are all first reference values, and the element values of the remaining elements are all second reference values, the first reference value indicates that the attention distribution is allowed to be calculated, and the second reference value indicates that the attention distribution is prohibited to be calculated.

[0101] The information extraction apparatus in the embodiments of the present application can be an apparatus or a component in an electronic device, such as an integrated circuit or a chip. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, and the embodiments of the present application are not limited in this regard.

[0102] The electronic device in the embodiments of the present application can be a terminal having an operating system. The operating system can be an Android operating system, can be an iOS operating system, or can be another possible operating system, and the embodiments of the present application are not limited in this regard.

[0103] The information extraction apparatus provided in the embodiments of the present application can implement each process in the information extraction method embodiment and achieve the same technical effects, and thus repeated description is omitted. Figures 1 to 9

[0104] As shown in Figure 11 The embodiments of the present application also provide an electronic device 1100, which includes a processor 1101 and a memory 1102. The memory 1102 stores programs or instructions executable on the processor 1101. When the programs or instructions are executed by the processor 1101, each step of the above information extraction method embodiment is implemented, and the same technical effects can be achieved. To avoid repetition, further description is omitted.

[0105] It should be noted that the electronic device in the embodiments of the present application includes the mobile terminal and the non-mobile terminal described above.

[0106] Figure 12 A hardware structure schematic diagram of an electronic device provided in the embodiments of the present application is shown in FIG. 10.

[0107] ​The electronic device 1200 includes, but is not limited to, a radio frequency unit 1201, a network module 1202, an audio output unit 1203, an input unit 1204, a sensor 1205, a display unit 1206, a user input unit 1207, an interface unit 1208, a memory 1209, and a processor 1210, etc.

[0108] Those skilled in the art can understand that the electronic device 1200 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1210 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system. Figure 12 The structure of the electronic device 1200 shown in the figure does not constitute a limitation on the electronic device 1200, and the electronic device 1200 can include more or fewer components than shown, or combine certain components, or different component arrangements, which are not described here.

[0109] The processor 1210 is configured to: obtain a configuration file corresponding to an image category of a first image, the first image including a plurality of information fields, the configuration file including a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image; process the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word being used to prompt extraction of field information of one information field; output the field information of each information field according to the sub-task prompt words.

[0110] The embodiments of the present application can obtain a pre-set configuration file according to the image category of the first image, determine the task prompt word corresponding to the first image based on the configuration file, and obtain the sub-task prompt word corresponding to each information field based on the task prompt word, thereby realizing splitting of one task of extracting a plurality of information fields into a plurality of sub-tasks, and parallel prediction of the field information of each information field based on the sub-task prompt words of the plurality of sub-tasks and the attention mask matrix, thereby improving the information extraction efficiency.

[0111] In some possible implementations of the embodiments of the present application, the processor 1210 is specifically configured to: splice the sub-task prompt words to obtain a spliced prompt word; input a token sequence corresponding to the spliced prompt word and the attention mask matrix into an information prediction model, and output each information field and the field information of each information field; The attention mask matrix is used to isolate the input sub-task prompt words corresponding to each information field and isolate the output field information of each information field.

[0112] In some possible implementations of the embodiments of the present application, the task prompt word includes a basic prompt word and a field prompt word. The processor 1210 is specifically configured to: extract the basic prompt word and the field prompt word of each information field from the task prompt word respectively; generate a sub-task prompt word corresponding to each information field according to the basic prompt word and the field prompt word of each information field.

[0113] In some possible implementations of the embodiments of the present application, the attention mask matrix includes a first attention mask matrix and a second attention mask matrix, the first attention mask matrix is used for isolating the sub-task prompt word corresponding to each information field of the input, and the second attention mask matrix is used for isolating the field information of each information field of the output; The processor 1210 is specifically configured to: input the token sequence corresponding to the spliced prompt word and the first attention mask matrix into a first information prediction module of the information prediction model, and output a first token sequence, the first token sequence including one output token corresponding to each information field; input the second token sequence, the token sequence corresponding to the spliced prompt word and the second attention mask matrix into a second information prediction module of the information prediction model, and output a third token sequence; the token corresponding to each information field in the third token sequence is a next token of the token corresponding to each information field in the second token sequence, the second token sequence is the output token corresponding to each information field that has been predicted, and the second token sequence includes the first token sequence; output each information field and the field information of each information field according to all the token sequences predicted by the information prediction model.

[0114] In some possible implementations of the embodiments of the present application, the attention mask matrix includes a first attention mask matrix, and the first attention mask matrix is used for isolating the sub-task prompt word corresponding to each information field of the input; The processor 1210 is further configured to, after obtaining the configuration file corresponding to the image category of the first image, determine the number of information fields contained in the first image according to the configuration file, and determine the token corresponding to each field prompt word according to the field prompt word in the task prompt word. generate the first attention mask matrix according to the number of information fields and the token corresponding to each field prompt word, the number of rows and the number of columns of the first attention mask matrix being the same as the total number of tokens corresponding to the field prompt words; An element value of each element in the first attention mask matrix is determined according to a position of the element and an information field to which a token corresponding to the element belongs; In the first attention mask matrix, element values of lower triangular elements of the token matrix corresponding to a same information field are all first reference values, and element values of the remaining elements are all second reference values. The first reference value indicates that the attention distribution is allowed to be calculated, and the second reference value indicates that the attention distribution is prohibited to be calculated.

[0115] It should be understood that in the embodiments of the present application, the input unit 1204 can include a graphics processing unit (GPU) 12041 and a microphone 12042. The graphics processing unit 12041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1206 can include a display panel 12061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1207 includes at least one of a touch panel 12071 and other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 can include two parts of a touch detection device and a touch controller. The other input devices 12072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.

[0116] The memory 1209 can be used to store software programs and various data. The memory 1209 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 1209 can include a volatile memory or a non-volatile memory, or the memory 1209 can include both volatile and non-volatile memories. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1209 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0117] The processor 1210 can include one or more processing units; optionally, the processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1210.

[0118] The embodiments of the present application also provide a readable storage medium, and the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize each process of the above-mentioned information extraction method embodiments, and the same technical effects can be achieved, and thus details are not repeated here.

[0119] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0120] The chip provided in the embodiments of the present application includes a processor and a communication interface, the communication interface is coupled with the processor, the processor is used to run programs or instructions to realize the processes of the above information extraction method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0121] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip, etc.

[0122] The embodiments of the present application provide a computer program product stored in a storage medium, which is executed by at least one processor to realize the processes of the above information extraction method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0123] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article, or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in a reverse order, for example, the described method can be performed in an order different from the described order, and various steps can be added, omitted, or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0124] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in the embodiments of the present application.

[0125] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.

Claims

1. An information extraction method characterized by, The method comprises the following steps: obtaining a configuration file corresponding to an image category of a first image, the first image comprising a plurality of information fields, the configuration file comprising a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image; processing the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word being used to prompt extraction of field information of one information field; outputting the field information of each information field according to the sub-task prompt words.

2. The method of claim 1, wherein, The step of outputting the field information of each information field according to the sub-task prompt words comprises the following steps: splicing the sub-task prompt words to obtain a spliced prompt word; inputting a token sequence corresponding to the spliced prompt word and an attention mask matrix into an information prediction model to output each information field and the field information of each information field; wherein the attention mask matrix is used to isolate the input sub-task prompt words corresponding to each information field and isolate the output field information of each information field.

3. The method of claim 2, wherein, The task prompt word comprises a basic prompt word and a field prompt word; The step of processing the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field comprises the following steps: extracting the basic prompt word and the field prompt word of each information field from the task prompt word; generating a sub-task prompt word corresponding to each information field according to the basic prompt word and the field prompt word of each information field.

4. The method of claim 2, wherein, The attention mask matrix comprises a first attention mask matrix and a second attention mask matrix, the first attention mask matrix being used to isolate the input sub-task prompt words corresponding to each information field, and the second attention mask matrix being used to isolate the output field information of each information field; The step of inputting the token sequence corresponding to the spliced prompt word and the attention mask matrix into the information prediction model to output each information field and the field information of each information field comprises the following steps: inputting the token sequence corresponding to the spliced prompt word and the first attention mask matrix into a first information prediction module of the information prediction model to output a first token sequence, the first token sequence comprising one output token corresponding to each information field; inputting a second token sequence, the token sequence corresponding to the spliced prompt word and the second attention mask matrix into a second information prediction module of the information prediction model to output a third token sequence; the tokens corresponding to each information field in the third token sequence are the next tokens of the tokens corresponding to each information field in the second token sequence, the second token sequence being the output tokens corresponding to each information field that have been predicted, and the second token sequence comprising the first token sequence; outputting each information field and the field information of each information field according to all token sequences predicted by the information prediction model.

5. The method according to any one of claims 2-4, characterized in that, The attention mask matrix comprises a first attention mask matrix, and the first attention mask matrix is used for isolating the sub-task prompt words corresponding to the information fields of the input; After the configuration file corresponding to the image category of the first image is obtained, the method further comprises: determining the number of information fields contained in the first image according to the configuration file, and determining the token corresponding to each field prompt word according to the field prompt words in the task prompt word; generating a first attention mask matrix according to the number of information fields and the token corresponding to each field prompt word, wherein the number of rows and the number of columns of the first attention mask matrix are the same as the total number of tokens corresponding to each field prompt word; determining the element value of each element in the first attention mask matrix according to the position of each element and the information field to which the element corresponds; wherein the element values of the lower triangular elements of the token matrix corresponding to the same information field in the first attention mask matrix are all first reference values, and the element values of the remaining elements are all second reference values, the first reference value indicating that the attention distribution is allowed to be calculated, and the second reference value indicating that the attention distribution is prohibited to be calculated.

6. An information extraction apparatus characterized by comprising: comprises: an acquisition module configured to acquire a configuration file corresponding to an image category of a first image, the first image comprising a plurality of information fields, and the configuration file comprising a task prompt word, the task prompt word being used to prompt extraction of field information of all information fields in the first image; a processing module configured to process the task prompt word according to the information fields to obtain a sub-task prompt word corresponding to each information field, the sub-task prompt word being used to prompt extraction of field information of one information field; an output module configured to output field information of each information field according to each sub-task prompt word.

7. The apparatus of claim 6, wherein, The output module is specifically configured to: splice each sub-task prompt word to obtain a spliced prompt word; input a token sequence corresponding to the spliced prompt word and an attention mask matrix into an information prediction model to output each information field and field information of each information field; wherein the attention mask matrix is used for isolating the sub-task prompt words corresponding to the input information fields and isolating the field information of the output information fields.

8. The apparatus of claim 7, wherein, The task prompt word comprises a basic prompt word and a field prompt word; The processing module is specifically configured to: extract the basic prompt word and the field prompt word of each information field from the task prompt word; generate a sub-task prompt word corresponding to each information field according to the basic prompt word and the field prompt word of each information field.

9. The apparatus of claim 7, wherein, The attention mask matrix comprises a first attention mask matrix and a second attention mask matrix, the first attention mask matrix is used for isolating the sub-task prompt words corresponding to the input information fields, and the second attention mask matrix is used for isolating the field information of the output information fields; The output module is specifically configured to: The token sequence corresponding to the spliced prompt word and the first attention mask matrix are input into a first information prediction module of the information prediction model, and a first token sequence is output, the first token sequence including one output token corresponding to each information field; The second token sequence, the token sequence corresponding to the spliced prompt word, and the second attention mask matrix are input into a second information prediction module of the information prediction model, and a third token sequence is output; the token corresponding to each information field in the third token sequence is the next token of the token corresponding to each information field in the second token sequence, the second token sequence being the output token corresponding to each information field that has been predicted, and the second token sequence including the first token sequence; According to all token sequences predicted by the information prediction model, each information field and field information of each information field are output.

10. The device of any of claims 7-9, wherein, The attention mask matrix includes a first attention mask matrix, and the first attention mask matrix is used to isolate the input subtask prompt word corresponding to each information field; The processing module is further configured to, after the acquisition module acquires the configuration file corresponding to the image category of the first image, determine the number of information fields contained in the first image according to the configuration file, and determine the token corresponding to each field prompt word according to the field prompt word in the task prompt word; According to the number of information fields and the token corresponding to each field prompt word, a first attention mask matrix is generated, the number of rows and the number of columns of the first attention mask matrix being the same as the total number of tokens corresponding to each field prompt word; According to the position of each element in the first attention mask matrix and the information field to which the token corresponding to the element belongs, an element value of each element is determined; In the first attention mask matrix, the element values of the lower triangular elements of the token matrix corresponding to the same information field are all first reference values, and the element values of the remaining elements are all second reference values, the first reference value indicating that the attention distribution is allowed to be calculated, and the second reference value indicating that the attention distribution is prohibited to be calculated.

11. An electronic device, comprising: The electronic device includes a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the method of any one of claims 1-5.