Text acquisition method and related device
By combining image features and text location information with a target model and utilizing neural network decoding technology, the problem of inaccurate text extraction by pre-trained models was solved, achieving accurate text acquisition and visual interaction, and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2023-04-28
- Publication Date
- 2026-04-17
AI Technical Summary
Existing pre-trained neural network models consider only a few factors when extracting target text from images, resulting in inaccurate extracted text and an inability to provide visual interaction or output longer text, thus reducing the user experience.
The target image is encoded by the target model to obtain image features. Combined with the position information of the target text in the image, the target text is decoded using recurrent neural networks, multilayer perceptrons or temporal convolutional networks to obtain accurate target text, which can be converted into visual coordinates or character form.
It achieves a full and accurate understanding of the target image content, can extract the correct target text, and provides visualization effects, thus improving the user experience.
Smart Images

Figure CN116758572B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI), and more particularly to a text acquisition method and related equipment. Background Technology
[0002] With the rapid development of AI technology, more and more users are using pre-trained neural network models (also known as pre-trained models) to perform analysis and processing of images that present multiple texts. In other words, pre-trained neural network models can fully understand images to extract target text from the multiple texts presented in the image.
[0003] In related technologies, pre-trained neural network models can include encoders and decoders. When it is necessary to extract target text from multiple texts presented in an image, the image can be input into the neural network model. The encoder can then encode the image to obtain its features and provide these features to the decoder. The decoder can then decode based on these features to obtain the target text.
[0004] In the above process, the neural network model understands the content of an image based on its features in order to extract the target text from the multiple texts presented in the image. However, the neural network model considers relatively few factors when understanding an image, which may result in the target text obtained by the model being inaccurate. Summary of the Invention
[0005] This application provides a text acquisition method and related equipment, which can acquire accurate target text from a target image.
[0006] A first aspect of this application provides a text acquisition method, which is implemented through a target model and includes:
[0007] When you need to extract target text from a target image, you can first extract the target image. It should be noted that the content presented in the target image contains multiple texts, and these multiple texts contain the target text that you want to extract.
[0008] After obtaining the target image, it can be input into the target model. The target model can then encode the target image to obtain its features. After obtaining these features, the target model can process them to obtain the location information of the target text within the target image. Finally, after obtaining the location information of the target text, the decoder can further process the target image features and the location information of the target text to obtain the target text.
[0009] It should be noted that the input to the target model includes not only the externally input target image, but also the positional information of the target text within the target image. The output of the target model includes not only the target text, but also the positional information of the target text within the target image. In other words, the target text and its positional information within the target image are the two outputs of the target model. Thus, the target text has been successfully extracted from the target image.
[0010] As can be seen from the above method, when extracting target text from a target image, the first step is to acquire a target image containing multiple texts and input it into a target model. Next, the target model encodes the target image to obtain its features. Then, the target model processes these features to obtain the positional information of the target text within the target image. Finally, the target model further processes the features and positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, the target model considers not only the features of the target image but also the positional information of the target text when understanding its content. This comprehensive approach allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted from multiple texts presented in the target image using this method is typically the correct text.
[0011] In one possible implementation, obtaining the location information of the target text in the target image based on features includes: decoding the first to the ith vector representation of the location information of the target text in the target image based on features to obtain the (i+1)th vector representation of the location information, where i = 1, ..., X-1, X ≥ 1. The first vector representation of the location information is obtained by decoding a preset vector representation based on features. In the aforementioned implementation, if there is only one target text, after obtaining the features of the target image, the target model can first decode the preset vector representation based on the features of the target image to obtain the first vector representation of the location information of the target text in the target image. Next, the target model can decode the first vector representation of the target text's position information in the target image based on the features of the target image, thereby obtaining the second vector representation of the target text's position information in the target image, and so on. Finally, the target model can decode the first to the (X-1)th vector representations of the target text's position information in the target image based on the features of the target image, thereby obtaining the Xth vector representation of the target text's position information in the target image. In this way, the target model can accurately obtain the position information of the target text in the target image, presented in vector representation form.
[0012] In one possible implementation, obtaining the target text based on features and location information includes: decoding the first to the j-th vector representation of the target text based on the features to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on the features. In the aforementioned implementation, if there is only one target text, after obtaining the location information of the target text in the target image, the target model can first decode the location information of the target text in the target image based on the features of the target image to obtain the first vector representation of the target text. Next, the target model can decode the location information of the target text in the target image and the first vector representation of the target text based on the features of the target image to obtain the second vector representation of the target text, ..., and finally, the target model can decode the location information of the target text in the target image and the first to the (Y-1)-th vector representations of the target text based on the features of the target image to obtain the Y-th vector representation of the target text. In this way, the target model can accurately obtain the target text presented in vector representation.
[0013] In one possible implementation, the target text includes a first text and a second text. The location information includes the first location information of the first text in the target image and the second location information of the second text in the target image. Obtaining the location information of the target text in the target image based on features includes: decoding the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, where i = 1, ..., X-1, X ≥ 1; and decoding the first vector representation of the first location information based on features using a preset vector representation. Furthermore, decoding the first vector representation to the k-th vector representation of the second location information, the first text, and the first vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, where k = 1, ..., Z-1, Z ≥ 1; and decoding the first vector representation of the second location information based on features using the first location information and the first text. In the aforementioned implementation, if there are two target texts, these two target texts can be referred to as the first text and the second text, respectively. After obtaining the features of the target image, the target model can first decode the preset vector representation based on the features of the target image to obtain the first vector representation of the first position information of the first text in the target image. Next, the target model can decode the first vector representation of the first position information based on the features of the target image to obtain the second vector representation of the first position information, and so on. Finally, the target model can decode the first to the (X-1)th vector representations of the first position information based on the features of the target image to obtain the Xth vector representation of the first position information. In this way, the target model can obtain the complete first position information presented in vector representation form. After obtaining the first position information, the target model can process the first position information based on the features of the target image to obtain the first text. After obtaining the first text, the target model can first decode the first position information and the first text based on the features of the target image to obtain the first vector representation of the second position information of the second text in the target image. Next, the target model can decode the first location information, the first text, and the first vector representation of the second location information based on the features of the target image, thereby obtaining the second vector representation of the second location information, and so on. Finally, the target model can decode the first location information, the first text, and the first to the (Z-1)th vector representations of the second location information based on the features of the target image, thereby obtaining the Zth vector representation of the second location information. In this way, the target model can accurately obtain the second location information presented in vector representation form.
[0014] In one possible implementation, obtaining the target text based on features and location information includes: decoding the first location information and the first vector representation to the j-th vector representation of the first text based on features to obtain the (j+1)-th vector representation of the first text, where j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the location information based on features; and decoding the first location information, the first text, the second location information, and the first vector representation to the t-th vector representation of the second text based on features to obtain the (t+1)-th vector representation of the second text, where t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first location information, the first text, and the second location information based on features. In the aforementioned implementation, if there are two target texts, these two target texts can be referred to as the first text and the second text, respectively. After obtaining the first location information of the first text in the target image, the target model can first decode the first location information based on the features of the target image to obtain the first vector representation of the first text. Next, the target model can decode the first position information and the first vector representation of the first text based on the features of the target image, thereby obtaining the second vector representation of the first text, and so on. Finally, the target model can decode the first position information and the first to (Y-1)th vector representations of the first text based on the features of the target image, thereby obtaining the Yth vector representation of the first text. In this way, the target model can obtain the first text presented in vector representation form. After obtaining the first text, the target model can process the first position information and the first text based on the features of the target image, thereby obtaining the second position information of the second text in the target image. After obtaining the second position information of the second text in the target image, the target model can first decode the first position information, the first text, and the second position information based on the features of the target image, thereby obtaining the first vector representation of the second text. Next, the target model can decode the first positional information, the first text, the second positional information, and the first vector representation of the second text based on the features of the target image, thereby obtaining the second vector representation of the second text, and so on. Finally, the target model can decode the first positional information, the first text, the second positional information, and the first to (U-1)th vector representations of the second text based on the features of the target image, thereby obtaining the U-th vector representation of the second text. In this way, the target model can accurately obtain the second text presented in vector representation form.
[0015] In one possible implementation, the method further includes: transforming all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image. In the aforementioned implementation, the target model can convert the location information of the target text (in the target image) presented in vector representation into the location information of the target text presented in coordinate form, thereby providing users with a visualization of the target text in the target image.
[0016] In one possible implementation, the method further includes: transforming all vector representations of the target text to obtain all characters of the target text. In the aforementioned implementation, the target model can also convert the target text presented in vector representation into target text presented in character (text) form, further providing users with a visualization of the target text within the target image.
[0017] In one possible implementation, the transformation performed by the target model on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0018] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0019] A second aspect of this application provides a model training method, the method comprising: acquiring a target image, the target image containing multiple texts; processing the target image using a model to be trained to obtain the location information of the target text in the target image and the target text, wherein the multiple texts contain the target text, and the model to be trained is used to: encode the target image to obtain features of the target image; acquire location information based on the features; acquire the target text based on the features and the location information; and train the model to be trained based on the target text to obtain a target model.
[0020] The target text trained using the above method possesses text extraction capabilities. Specifically, when extracting target text from a target image, a target image containing multiple texts can be acquired first and input into the target model. Next, the target model encodes the target image to obtain its features. Then, the target model processes these features to obtain the positional information of the target text within the target image. Finally, the target model further processes the target image's features and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the image's features but also the positional information of the target text within the image. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from multiple texts presented in the target image using this method is typically the correct text.
[0021] In one possible implementation, the model to be trained is used to decode the first to the ith vector representation of the positional information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the positional information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the positional information is obtained by decoding the preset vector representation based on features.
[0022] In one possible implementation, the model to be trained is used to decode the first to the j-th vector representation of the target text based on the feature-based location information to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on the feature.
[0023] In one possible implementation, the target text includes a first text and a second text, and the location information includes the first location information of the first text in the target image and the second location information of the second text in the target image. The model to be trained is used to: decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, and the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, and the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0024] In one possible implementation, the model to be trained is used to: decode the first positional information, the first vector representation of the first text to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the positional information based on features; and decode the first positional information, the first text, the second positional information, the first vector representation of the second text to the t-th vector representation of the second text based on features, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first positional information, the first text, and the second positional information based on features.
[0025] In one possible implementation, the model to be trained is also used to transform all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
[0026] In one possible implementation, the model to be trained is also used to transform all vector representations of the target text to obtain all characters of the target text. Training the model to be trained based on the target text to obtain the target model includes: training the model to be trained based on characters and coordinates to obtain the target model.
[0027] In one possible implementation, the transformation performed by the model to be trained on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0028] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0029] A third aspect of this application provides a text acquisition device, which includes a target model. The device includes: a first acquisition module for acquiring a target image, the target image containing multiple texts; an encoding module for encoding the target image to obtain features of the target image; a second acquisition module for acquiring positional information of target text in the target image based on the features, the multiple texts containing the target text; and a third acquisition module for acquiring the target text based on the features and the positional information.
[0030] As can be seen from the above apparatus, when it is necessary to extract target text from a target image, a target image containing multiple texts can be acquired first and input into a target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process these features to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within it. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from multiple texts presented in the target image using this method is usually the correct text.
[0031] In one possible implementation, the second acquisition module is used to decode the first vector representation to the ith vector representation of the position information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the position information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the position information is obtained by decoding the preset vector representation based on features.
[0032] In one possible implementation, the third acquisition module is used to decode the location information, from the first vector representation to the j-th vector representation of the target text, based on features, to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on features.
[0033] In one possible implementation, the target text includes first text and second text, and the location information includes first location information of the first text in the target image and second location information of the second text in the target image. The second acquisition module is used to decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, where the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and to decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, where the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0034] In one possible implementation, the third acquisition module is used to decode the first position information and the first vector representation of the first text up to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, where the first vector representation of the first text is obtained by decoding the position information based on features; and to decode the first position information, the first text, the second position information, and the first vector representation of the second text up to the t-th vector representation of the second text based on features, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, where the first vector representation of the second text is obtained by decoding the first position information, the first text, and the second position information based on features.
[0035] In one possible implementation, the device further includes: a first transformation module for transforming all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
[0036] In one possible implementation, the device further includes a second conversion module for converting all vector representations of the target text to obtain all characters of the target text.
[0037] In one possible implementation, the transformation performed by the target model on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0038] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0039] A fourth aspect of this application provides a model training apparatus, comprising: an acquisition module for acquiring a target image containing multiple texts; a processing module for processing the target image using a model to be trained to obtain location information of the target text in the target image and the target text itself, wherein the multiple texts contain the target text; the model to be trained being configured to: encode the target image to obtain features of the target image; acquire location information based on the features; acquire the target text based on the features and location information; and a training module for training the model to be trained based on the target text to obtain a target model.
[0040] The target text trained by the aforementioned device possesses text acquisition capabilities. Specifically, when it is necessary to extract target text from a target image, a target image containing multiple texts can first be acquired and input into the target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process these features to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within it. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from multiple texts presented in the target image using this method is typically the correct text.
[0041] In one possible implementation, the model to be trained is used to decode the first to the ith vector representation of the positional information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the positional information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the positional information is obtained by decoding the preset vector representation based on features.
[0042] In one possible implementation, the model to be trained is used to decode the first to the j-th vector representation of the target text based on the feature-based location information to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on the feature.
[0043] In one possible implementation, the target text includes a first text and a second text, and the location information includes the first location information of the first text in the target image and the second location information of the second text in the target image. The model to be trained is used to: decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, and the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, and the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0044] In one possible implementation, the model to be trained is used to: decode the first positional information, the first vector representation of the first text to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the positional information based on features; and decode the first positional information, the first text, the second positional information, the first vector representation of the second text to the t-th vector representation of the second text based on features, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first positional information, the first text, and the second positional information based on features.
[0045] In one possible implementation, the model to be trained is also used to transform all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
[0046] In one possible implementation, the model to be trained is also used to transform all vector representations of the target text to obtain all characters of the target text. The training module is used to train the model to be trained based on the characters and coordinates to obtain the target model.
[0047] In one possible implementation, the transformation performed by the model to be trained on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0048] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0049] A fifth aspect of this application provides a fault prediction apparatus, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the fault prediction apparatus performs the method described in the first aspect or any possible implementation thereof.
[0050] A sixth aspect of this application provides a model training apparatus, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training apparatus performs the method described in the second aspect or any possible implementation thereof.
[0051] A seventh aspect of this application provides a circuit system including a processing circuit configured to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0052] An eighth aspect of this application provides a chip system including a processor for calling a computer program or computer instructions stored in a memory, such that the processor performs the method as described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0053] In one possible implementation, the processor is coupled to the memory via an interface.
[0054] In one possible implementation, the chip system also includes a memory that stores computer programs or computer instructions.
[0055] A ninth aspect of this application provides a computer storage medium storing a computer program that, when executed by a computer, causes the computer to perform the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0056] A tenth aspect of this application provides a computer program product storing instructions that, when executed by a computer, cause the computer to perform the method as described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.
[0057] In this embodiment, when it is necessary to extract target text from a target image, a target image containing multiple texts can first be obtained and input into a target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process these features to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within it. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from the multiple texts presented in the target image in this manner is usually the correct text. Attached Figure Description
[0058] Figure 1 A structural diagram illustrating the main framework of artificial intelligence;
[0059] Figure 2a A schematic diagram of the structure of the text acquisition system provided in the embodiments of this application;
[0060] Figure 2b Another structural schematic diagram of the text acquisition system provided in the embodiments of this application;
[0061] Figure 2c A schematic diagram of a text acquisition device provided in an embodiment of this application;
[0062] Figure 3 A schematic diagram of the system 100 architecture provided in the embodiments of this application;
[0063] Figure 4A schematic diagram of the structure of the target model provided in the embodiments of this application;
[0064] Figure 5 A flowchart illustrating the text acquisition method provided in this application embodiment;
[0065] Figure 6 Another structural schematic diagram of the target model provided in the embodiments of this application;
[0066] Figure 7 Another structural schematic diagram of the target model provided in the embodiments of this application;
[0067] Figure 8 Another structural schematic diagram of the target model provided in the embodiments of this application;
[0068] Figure 9 A schematic diagram of the document Q&A provided for embodiments of this application;
[0069] Figure 10 Another structural schematic diagram of the target model provided in the embodiments of this application;
[0070] Figure 11 A schematic diagram illustrating information extraction provided in an embodiment of this application;
[0071] Figure 12 A schematic flowchart of the model training method provided in the embodiments of this application;
[0072] Figure 13 A schematic diagram of the structure of the text acquisition device provided in the embodiments of this application;
[0073] Figure 14 A schematic diagram of the structure of the model training apparatus provided in the embodiments of this application;
[0074] Figure 15 A schematic diagram of the structure of the execution device provided in the embodiments of this application;
[0075] Figure 16 A schematic diagram of the structure of the training device provided in the embodiments of this application;
[0076] Figure 17 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation
[0077] This application provides a text acquisition method and related equipment, which can acquire accurate target text from a target image.
[0078] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0079] With the rapid development of AI technology, more and more users are using pre-trained neural network models (also known as pre-trained models) to perform analysis and processing of images that present multiple texts. In other words, pre-trained neural network models can fully understand images to extract target text from the multiple texts presented in the image.
[0080] In related technologies, pre-trained neural network models can include encoders and decoders. When a user needs to extract target text from multiple texts presented in an image, the image can be input into the neural network model. The encoder encodes the image to obtain its features and provides these features to the decoder. The decoder then decodes the image based on these features to obtain and return the target text to the user. For example, when a user needs to extract the passenger's name from an image of a train ticket, the image can be input into a pre-trained model. The pre-trained model can extract the image's features and, based on these features, extract the text "passenger's name" from multiple texts presented in the image, such as "passenger's name," "train number," "time," "departure point," and "destination," and return it to the user.
[0081] In the above process, the neural network model understands the content of an image based on its features in order to extract the target text from the multiple texts presented in the image. However, the neural network model considers relatively few factors when understanding an image, which may result in the target text obtained by the model being inaccurate.
[0082] Furthermore, in the above process, neural network models typically only output the target text to the user, failing to provide a reasonable explanation of the output (i.e., unable to explain why the model extracted the target text) and visual interaction (i.e., unable to provide any additional content related to the target text besides the text itself), thus reducing the user experience.
[0083] Furthermore, in the above process, the output length of the neural network model is limited, and in some special scenarios, users often need to obtain longer texts, which the model cannot meet, further reducing the user experience.
[0084] To address the aforementioned problems, this application provides a text acquisition method that can be implemented in conjunction with artificial intelligence (AI) technology. AI technology is a discipline that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence. AI technology achieves optimal results by perceiving the environment, acquiring knowledge, and using that knowledge. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. Using artificial intelligence for data processing is a common application of AI.
[0085] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 This is a structural diagram illustrating the main framework of artificial intelligence. The following explanation of the AI framework is based on two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.
[0086] (1) Infrastructure
[0087] Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips (hardware acceleration chips such as CPUs, NPUs, GPUs, ASICs, and FPGAs); and the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0088] (2) Data
[0089] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, as well as IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0090] (3) Data processing
[0091] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0092] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data by symbolizing and formalizing it.
[0093] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0094] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0095] (4) General ability
[0096] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0097] (5) Smart Products and Industry Applications
[0098] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0099] The following sections will introduce several application scenarios for this application.
[0100] Figure 2a This is a schematic diagram of a text acquisition system provided in an embodiment of this application. The text acquisition system includes a user device and a data processing device. The user device includes smart terminals such as mobile phones, personal computers, or information processing centers. The user device is the initiator of the text acquisition request; typically, the request is initiated by the user through the user device.
[0101] The aforementioned data processing equipment can be cloud servers, network servers, application servers, management servers, or other devices or servers with data processing capabilities. The data processing equipment receives text processing requests from smart terminals through an interactive interface, and then performs text processing through a storage device for storing data and a processor for data processing, employing methods such as machine learning, deep learning, search, reasoning, and decision-making. The storage device in the data processing equipment can be a general term, including local storage and a database storing historical data. The database can be located on the data processing equipment or on other network servers.
[0102] exist Figure 2a In the text acquisition system shown, the user device can receive user instructions. For example, the user device can acquire a target image containing multiple texts, input / selected by the user, and then send a request to the data processing device. This causes the data processing device to perform image processing on the target image obtained by the user device, thereby obtaining the corresponding processing result for the image. For instance, the user device can acquire a target image input by the user (the target image contains multiple texts), and then send a processing request to the data processing device. This causes the data processing device to perform a series of processes on the target image, thereby obtaining the processing result of the target image, namely, the target text among the multiple texts contained in the target image and the position information of the target text within the target image.
[0103] exist Figure 2a In this context, the data processing device can execute the text acquisition method of the embodiments of this application.
[0104] Figure 2b This is another schematic diagram of the structure of the text acquisition system provided in the embodiments of this application. Figure 2b In this context, the user equipment (UE) directly functions as a data processing device. This UE can directly acquire input from the user and process it directly through its own hardware. The specific process is similar to... Figure 2a Similar to the description above, it will not be repeated here.
[0105] exist Figure 2b In the text acquisition system shown, the user equipment can receive user instructions. For example, the user equipment can acquire a target image input by the user (the content presented in the target image contains multiple texts), and then perform a series of processing on the target image to obtain the processing result of the target image, namely the target text in the multiple texts contained in the target image and the position information of the target text in the target image.
[0106] exist Figure 2b In this context, the user equipment itself can execute the text acquisition method of this application embodiment.
[0107] Figure 2c This is a schematic diagram of a text acquisition device provided in an embodiment of this application.
[0108] The above Figure 2a and Figure 2b The user equipment in the context can specifically be Figure 2c Local device 301 or local device 302 in the system. Figure 2a The data processing equipment in the middle can specifically be Figure 2c The execution device 210 in the process includes a data storage system 250 that can store the data to be processed by the execution device 210. The data storage system 250 can be integrated into the execution device 210 or set up in the cloud or on other network servers.
[0109] Figure 2a and Figure 2b The processor in the image can be trained on data using neural network models or other models (e.g., support vector machine-based models) for machine learning / deep learning, and then use the trained or learned models to perform image processing applications on the image to obtain the corresponding processing results.
[0110] Figure 3 A schematic diagram of the system 100 architecture provided in this application embodiment, in Figure 3 In the process, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data in this embodiment may include various scheduled tasks, callable resources, and other parameters.
[0111] During the preprocessing of input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations and other related processing (such as implementing the neural network function in this application), the execution device 110 may call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0112] Finally, I / O interface 112 returns the processing result to client device 140, thereby providing it to the user.
[0113] It is worth noting that the training device 120 can generate corresponding target models / rules based on different training data for different objectives or tasks. These target models / rules can then be used to achieve the aforementioned objectives or complete the aforementioned tasks, thereby providing the user with the required results. The training data can be stored in the database 130 and originates from training samples collected by the data acquisition device 160.
[0114] exist Figure 3 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.
[0115] It is worth noting that, Figure 3 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 3 In this context, the data storage system 150 is an external memory relative to the execution device 110. However, in other cases, the data storage system 150 can also be placed within the execution device 110. For example... Figure 3 As shown, a neural network can be trained using training device 120.
[0116] This application also provides a chip including a neural network processor (NPU). This chip can be configured as follows: Figure 3 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be located in, for example... Figure 3 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rules.
[0117] The Neural Processing Unit (NPU) is a coprocessor mounted on the main central processing unit (CPU) (host CPU), where tasks are assigned by the CPU. The core of the NPU is the computation circuitry, which is controlled by a controller to retrieve data from memory (weight memory or input memory) and perform calculations.
[0118] In some implementations, the arithmetic circuitry includes multiple process engines (PEs). In some implementations, the arithmetic circuitry is a two-dimensional pulsating array. The arithmetic circuitry can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuitry is a general-purpose matrix processor.
[0119] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data for matrix B from the weight memory and caches it in each physical element (PE) of the arithmetic circuit. The arithmetic circuit retrieves the data for matrix A from the input memory and performs matrix operations with matrix B. The partial or final result of the obtained matrix is stored in the accumulator.
[0120] Vector computation units can further process the output of computational circuits, such as vector multiplication, vector addition, exponentiation, logarithmic operations, size comparisons, etc. For example, vector computation units can be used in non-convolutional / non-FC layers of neural networks for computation, such as pooling, batch normalization, and local response normalization.
[0121] In some implementations, the vector computation unit can store the processed output vector into a unified buffer. For example, the vector computation unit can apply a nonlinear function to the output of the arithmetic circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit, for example, for use in subsequent layers of a neural network.
[0122] The unified memory is used to store input data and output data.
[0123] The weight data is directly transferred from the external memory to the input memory and / or unified memory, stored in the weight memory, and stored in the unified memory to the external memory through the direct memory access controller (DMAC).
[0124] The bus interface unit (BIU) is used to enable interaction between the main CPU, DMAC, and instruction fetch memory via a bus.
[0125] The instruction fetch buffer, connected to the controller, is used to store the instructions used by the controller.
[0126] The controller is used to invoke instructions cached in the memory to control the operation of the computing accelerator.
[0127] Generally, the unified memory, input memory, weight memory, and instruction fetch memory are all on-chip memories, while external memory is memory outside the NPU. This external memory can be double data rate synchronous dynamic random access memory (DDRSDRAM), high bandwidth memory (HBM), or other readable and writable memories.
[0128] Since the embodiments of this application involve a large number of neural network applications, for ease of understanding, the relevant terms and concepts such as neural networks involved in the embodiments of this application will be introduced below.
[0129] (1) Neural Network
[0130] A neural network can be composed of neural units, which can be operational units that take xs and an intercept of 1 as inputs, and whose output can be:
[0131]
[0132] Where s = 1, 2, ..., n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.
[0133] The work of each layer in a neural network can be described by the mathematical expression y = a(Wx + b). From a physical perspective, the work of each layer in a neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are performed by Wx, operation 4 by +b, and operation 5 by a(). The term "space" is used here because the objects being classified are not individual things, but a class of things, and space refers to the set of all individuals of this class of things. Here, W is the weight vector, and each value in this vector represents the weight value of a neuron in that layer of the neural network. This vector W determines the spatial transformation from the input space to the output space mentioned above; that is, the weights W of each layer control how the space is transformed. The purpose of training a neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (a weight matrix formed by the vectors W of many layers). Therefore, the training process of a neural network is essentially about learning how to control the transformation space, and more specifically, learning the weight matrix.
[0134] Because we want the output of the neural network to be as close as possible to the actual predicted value, we can compare the current network's prediction with the desired target value, and then update the weight vector of each layer of the neural network based on the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring the parameters of each layer in the neural network). For example, if the network's prediction is too high, the weight vector is adjusted to make it predict lower, and this adjustment is continued until the neural network can predict the actual target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted value and the target value. Taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so training the neural network becomes the process of minimizing this loss as much as possible.
[0135] (2) Backpropagation algorithm
[0136] Neural networks can employ backpropagation (BP) to correct the parameters of the initial neural network model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates error loss; this error loss information is then propagated back to update the parameters of the initial neural network model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the neural network model, such as the weight matrix.
[0137] The method provided in this application is described below from the perspectives of neural network training and neural network application.
[0138] The model training method provided in this application involves the processing of data sequences and can be applied to data training, machine learning, deep learning, and other methods. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, and training on training data (e.g., the target image in the model training method of this application), ultimately obtaining a trained neural network (such as the target model in this application). Furthermore, the text acquisition method provided in this application can utilize the aforementioned trained neural network to input input data (e.g., the target image in the text acquisition method of this application) into the trained neural network, obtaining output data (e.g., the target text and its position information in the target image in the text acquisition method of this application). It should be noted that the model training method and text acquisition method provided in this application are inventions based on the same concept and can be understood as two parts of a system or two stages of a whole process: such as the model training stage and the model application stage.
[0139] The text acquisition method provided in this application embodiment can be implemented through a target model (also known as a pre-trained text (document) model). The structure of the target model will be briefly introduced below. Figure 4 A schematic diagram of the structure of the target model provided in the embodiments of this application, such as Figure 4 As shown, the input end of the target model can receive the target image from the outside and the position information of the target text in the target image from itself. The output end of the target model can output the target text and the position information of the target text in the target image. To understand... Figure 4 The workflow of the target model shown below, in conjunction with... Figure 5 This workflow will be described. Figure 5 A flowchart illustrating the text acquisition method provided in this application embodiment is shown below. Figure 5 As shown, the method includes:
[0140] 501. Obtain the target image, which contains multiple texts.
[0141] In this embodiment, when it is necessary to obtain target text from a target image, the target image can be obtained first. It should be noted that the content presented in the target image contains multiple texts. For example, when the target image is an income declaration form, the image contains multiple texts such as "Declarant's Name: Zhang XX", "Declaration Date: October 2, 20XX", "Income Amount: 200XXX Yuan", "Declarant's Gender: Male", and "Declarant's Contact Information: 130XXXXXXXX". Similarly, when the target image is a diabetes statistics report, the image contains multiple texts such as "Introduction to Diabetes: Diabetes is a common disease...", "Prevalence of Diabetes: 0.X", "Prevalence by Gender: 4.X% for Males and 3.X% for Females", "Annual Economic Losses Caused by Diabetes: 6X Billion Yuan", and "Preventability of Diabetes: 6X%".
[0142] It is understandable that the number of target texts to be obtained can be one or more; that is, it is necessary to obtain one or more texts contained in the target image. Continuing with the example above, if the target image is an income declaration form, the required target texts would be "Applicant Name: Zhang XX" and "Declaration Date: October 2, 20XX". Similarly, if the target image is a diabetes statistics report, the required target text would be "Preventable Probability of Diabetes: 6X%".
[0143] 502. Encode the target image to obtain its features.
[0144] After obtaining the target image, it can be input into the target model. The target model can first encode the target image to obtain its features.
[0145] Specifically, such as Figure 4 As shown, the target model can include an encoder and a decoder. Therefore, the target model can obtain the features of the target image in the following ways:
[0146] After the target image is input into the target model, the encoder of the target model can encode the target image to obtain the features of the target image, and then send the features of the target image to the decoder of the target model.
[0147] For example, such as Figure 6 As shown ( Figure 6(This is another structural diagram of the target model provided in an embodiment of this application). Suppose an image contains text 1, text 2, text 3, ..., text n (n is a positive integer greater than or equal to 2), and the target model includes an encoder and a decoder. When it is necessary to obtain text 1, text 2, ..., text m from the image (m is less than or equal to n, and m is a positive integer greater than or equal to 1), the image can be input into the target model. Then, after receiving the image, the encoder of the target model can encode the image to obtain its visual features. The encoding process is shown in the following formula:
[0148] H = Encoder(image) (2)
[0149] In the above formula, image represents the image itself, and H represents the visual features of the image. Therefore, after obtaining the visual features of the image, the encoder can provide these visual features to the decoder.
[0150] 503. Based on features, obtain the location information of the target text in the target image, where multiple texts contain the target text.
[0151] After obtaining the features of the target image, the target model can process the features of the target image to obtain the position information of the target text in the target image, and output the position information of the target text in the target image. In other words, the position information of the target text in the target image is one of the outputs of the target model.
[0152] Specifically, the target model can obtain the location information of the target text in the target image in several ways:
[0153] (1) If there is only one target text, after obtaining the features of the target image, the decoder can first decode the preset vector representation (also called the initial vector representation of the sequence, the content of which is usually meaningless) based on the features of the target image, thereby obtaining the first vector representation (token) of the target text's position information in the target image. Next, the decoder can decode the first vector representation of the target text's position information in the target image based on the features of the target image, thereby obtaining the second vector representation of the target text's position information in the target image. Subsequently, the decoder can decode the first and second vector representations of the target text's position information in the target image based on the features of the target image, thereby obtaining the third vector representation of the target text's position information in the target image,... Finally, the decoder can decode the first to the (X-1)th vector representations of the target text's position information in the target image based on the features of the target image, thereby obtaining the Xth vector representation of the target text's position information in the target image (X is a positive integer greater than or equal to 1). In this way, the decoder can obtain the location information of the target text in the target image, which is presented in vector form.
[0154] Continuing with the example above, suppose we only need to extract text 1 from the image. After obtaining the visual features of the image, the decoder can decode the beginning of sequence (BOS) based on these visual features, thereby obtaining the vector representation 1 of the position information 1 of text 1 (in the image), which is equivalent to obtaining the position information 1 presented in vector form.
[0155] (2) If there are two target texts, these two target texts can be referred to as the first text and the second text, respectively. After obtaining the features of the target image, the decoder can first decode the preset vector representation based on the features of the target image, thereby obtaining the first vector representation of the first position information of the first text in the target image. Next, the decoder can decode the first vector representation of the first position information based on the features of the target image, thereby obtaining the second vector representation of the first position information. Subsequently, the decoder can decode the first and second vector representations of the first position information based on the features of the target image, thereby obtaining the third vector representation of the first position information,... Finally, the decoder can decode the first to the (X-1)th vector representations of the first position information based on the features of the target image, thereby obtaining the Xth vector representation of the first position information. In this way, the decoder can obtain the complete first position information presented in vector representation form.
[0156] After obtaining the first location information, the decoder can process the first location information based on the features of the target image to obtain the first text. This process will not be elaborated on here.
[0157] After obtaining the first text, the decoder first decodes the first positional information and the first text based on the features of the target image, thereby obtaining the first vector representation of the second positional information of the second text in the target image. Next, the decoder decodes the first positional information, the first text, and the first vector representation of the second positional information based on the features of the target image, thereby obtaining the second vector representation of the second positional information. Subsequently, the decoder decodes the first positional information, the first text, the first vector representation of the second positional information, and the second vector representation of the second positional information based on the features of the target image, thereby obtaining the third vector representation of the second positional information, and so on. Finally, the decoder decodes the first positional information, the first text, and the first to the (Z-1)th vector representations of the second positional information based on the features of the target image, thereby obtaining the Zth vector representation of the second positional information (Z is a positive integer greater than or equal to 1). In this way, the decoder can obtain the second positional information presented in vector representation form.
[0158] Continuing with the example above, suppose we only need to extract text 1 and text 2 from the image. After obtaining the visual features of the image, the decoder can decode the initial vector representation of the sequence based on these visual features, thereby obtaining the vector representation 1 of the position information 1 of text 1 (in the image), which is equivalent to obtaining the position information 1 presented in vector representation form.
[0159] Next, the decoder can process the location information 1 based on the visual features of the image to obtain text 1, which will not be elaborated here.
[0160] Then, the decoder can also decode the vector representation 1 of position information 1, the vector representation 2 of text 1, the vector representation 3 of text 1 and the vector representation 4 of text 1 based on the visual features of the image, so as to obtain the vector representation 5 of position information 2 of text 2 (in the image), which is equivalent to obtaining position information 2 presented in vector representation form.
[0161] (3) If there are three or more target texts, that is, there are first text, second text and third text, etc., in this case, the process of the decoder obtaining the first position information of the first text in the target image, the second position information of the second text in the target image and the third position information of the third text in the target image is similar to the process described in (2) above, and will not be repeated here.
[0162] 504. Obtain the target text based on features and location information.
[0163] After obtaining the location information of the target text in the target image, the decoder can further process the features of the target image and the location information of the target text in the target image to obtain the target text and output the target text. In other words, the target text is another output of the target model.
[0164] Specifically, the target model can obtain the target text in several ways:
[0165] (1) If there is only one target text, after obtaining the position information of the target text in the target image, the decoder can first decode the position information of the target text in the target image based on the features of the target image, thereby obtaining the first vector representation of the target text. Next, the decoder can decode the position information of the target text in the target image and the first vector representation of the target text based on the features of the target image, thereby obtaining the second vector representation of the target text. Subsequently, the decoder can decode the position information of the target text in the target image and the first to second vector representations of the target text based on the features of the target image, thereby obtaining the third vector representation of the target text,... Finally, the decoder can decode the position information of the target text in the target image and the first to (Y-1)th vector representations of the target text based on the features of the target image, thereby obtaining the Yth vector representation of the target text (Y is a positive integer greater than or equal to 1). In this way, the decoder can obtain the target text presented in vector representation form.
[0166] Continuing with the example above, let's assume we only need to extract text 1 from the image. After obtaining the location information 1 of text 1, the decoder can decode the location information 1 based on the visual features of the image, thus obtaining the vector representation 2 of text 1. Next, the decoder can also decode the location information 1 and the vector representation 2 of text 1 based on the visual features of the image, thus obtaining the vector representation 3 of text 1. Then, the decoder can also decode the location information 1, the vector representation 2 of text 1, and the vector representation 3 of text 1 based on the visual features of the image, thus obtaining the vector representation 4 of text 1. At this point, it's equivalent to obtaining text 1 presented in vector form.
[0167] (2) If there are two target texts, these two target texts can be referred to as the first text and the second text, respectively. After obtaining the first position information of the first text in the target image, the decoder can first decode the first position information based on the features of the target image to obtain the first vector representation of the first text. Next, the decoder can decode the first position information and the first vector representation of the first text based on the features of the target image to obtain the second vector representation of the first text. Subsequently, the decoder can decode the first position information, the first vector representation of the first text to the second vector representation of the first text based on the features of the target image to obtain the third vector representation of the first text, ..., and finally, the decoder can decode the first position information, the first vector representation of the first text to the (Y-1)th vector representation of the first text based on the features of the target image to obtain the Yth vector representation of the first text. In this way, the decoder can obtain the first text presented in vector representation form.
[0168] After obtaining the first text, the decoder can process the first position information and the first text based on the features of the target image to obtain the second position information of the second text in the target image. This process can be referred to the relevant descriptions above, and will not be repeated here.
[0169] After obtaining the second position information of the second text in the target image, the decoder can first decode the first position information, the first text, and the second position information based on the features of the target image, thereby obtaining the first vector representation of the second text. Next, the decoder can decode the first position information, the first text, the second position information, and the first vector representation of the second text based on the features of the target image, thereby obtaining the second vector representation of the second text. Subsequently, the decoder can decode the first position information, the first text, the second position information, the first vector representation of the second text, and the second vector representation of the second text based on the features of the target image, thereby obtaining the third vector representation of the second text, and so on. Finally, the decoder can decode the first position information, the first text, the second position information, and the first to the (U-1)th vector representations of the second text based on the features of the target image, thereby obtaining the U-th vector representation of the second text (U is a positive integer greater than or equal to 1). In this way, the decoder can obtain the second text presented in vector representation form.
[0170] Continuing with the example above, let's assume we only need to extract text 1 from the image. After obtaining the location information 1 of text 1, the decoder can decode the location information 1 based on the visual features of the image, thus obtaining the vector representation 2 of text 1. Next, the decoder can also decode the location information 1 and the vector representation 2 of text 1 based on the visual features of the image, thus obtaining the vector representation 3 of text 1. Then, the decoder can also decode the location information 1, the vector representation 2 of text 1, and the vector representation 3 of text 1 based on the visual features of the image, thus obtaining the vector representation 4 of text 1. At this point, it's equivalent to obtaining text 1 presented in vector form.
[0171] Next, the decoder can process the location information 1 and text 1 based on the visual features of the image to obtain the location information 2 of text 2. This process can be referred to in the aforementioned relevant descriptions, and will not be repeated here.
[0172] Then, based on the visual features of the image, the decoder can decode the vector representation 1 of location information 1, the vector representation 2 of text 1, the vector representation 3 of text 1, the vector representation 4 of text 1, and the vector representation 5 of location information 2, thus obtaining the vector representation 6 of text 2. Next, based on the visual features of the image, the decoder can decode the vector representation 1 of location information 1, the vector representation 2 of text 1, the vector representation 3 of text 1, the vector representation 4 of text 1, the vector representation 5 of location information 2, and the vector representation 6 of text 2, thus obtaining the vector representation 7 of text 2. At this point, it is equivalent to obtaining text 2 presented in vector representation form.
[0173] (3) If the number of target texts is three or more, that is, there is a first text, a second text, a third text, etc., in this case, the process of the decoder to obtain the first text, the second text, the third text, etc. is similar to the process described in (2) above, and will not be repeated here.
[0174] More specifically, such as Figure 7 As shown ( Figure 7 (This is another structural schematic diagram of the target model provided in an embodiment of this application). In addition to including an encoder and a decoder, the target model may also include a converter. It should be noted that... Figure 4 In the target model shown, the decoder outputs the target text in vector representation and its position information within the target image. Figure 7In the target model shown, the decoder can send the target text, presented in vector representation, to the converter. The converter can transform all vector representations of the target text to obtain all characters of the target text, thus outputting the target text in character (text) form. Similarly, the decoder can also send the position information of the target text, presented in vector representation, within the target image to the conversion model. The conversion model can transform all vector representations of the target text's position information within the target image to obtain the coordinates of the area occupied by the target text in the target image, thus outputting the position information of the target text within the target image in coordinate form.
[0175] More specifically, in Figure 7 In the target model shown, the converter can be at least one of the following: a recurrent neural network, a multilayer perceptron, and a temporal convolutional network. Correspondingly, the transformation performed by the converter on the target text and location information can be at least one of the following: feature extraction based on a recurrent neural network, feature extraction based on a multilayer perceptron, and feature extraction based on a temporal convolutional network.
[0176] More specifically, the coordinates of the region occupied by the target text in an image typically include the coordinates of the top-left and bottom-right corners of that region. It should be noted that the conversion module can obtain the vertex coordinates of this region using the following formula:
[0177]
[0178] In the above formula, the conversion module can calculate all vector representations of the target text's position information in the target image, thereby obtaining the feature of the x-coordinate of the top-left vertex of the region. Again The calculation is performed to obtain the x-coordinate Z1 of the top-left vertex. Then, the conversion module can... And Z1 is used to calculate, thereby obtaining the characteristic ordinate of the top left corner vertex of the region. Again Calculations are performed to obtain the ordinate Z2 of the top-left vertex. This process is repeated until the x-coordinate Z1 of the top-left vertex, the ordinate Z2 of the top-left vertex, the x-coordinate Z3 of the bottom-right vertex, and the ordinate Z4 of the bottom-right vertex of the region are obtained.
[0179] It should be understood that in this embodiment, the coordinates of the region are only illustrated using the coordinates of the top-left and bottom-right vertices of the region. In practical applications, the coordinates of the region can also be any of the following: (1) the coordinates of the top-right and bottom-left vertices of the region; (2) the coordinates of the four corners of the region; (3) the coordinates of the top-left, bottom-left, and center points of the region; (4) the coordinates of the top-right, bottom-right, and center points of the region; (5) the coordinates of the top-right, top-left, and center points of the region; (6) the coordinates of the bottom-right, bottom-left, and center points of the region; (7) the coordinates of the top-right, bottom-right, top-left, bottom-left, and center points of the region, etc.
[0180] Furthermore, in Figure 4 or Figure 7 In the target model shown, the total output length of the decoder is unlimited. Generally, when the total number of required vector representations (positional information and text) is less than or equal to a preset threshold (e.g., 1024), the decoder can output all vector representations at once. When the total number of required vector representations is greater than the preset threshold, the decoder can output these vector representations in batches. For example, when 2000 vector representations are required, the decoder can first output the first to the 1024th vector representations as the first batch of output. Then, the decoder can use the last 25% of the vector representations from the first batch of output as input to continue decoding, thereby outputting the 1025th to the 2000th vector representations as the second batch of output. Therefore, the target model provided in this embodiment can ultimately output a sufficiently large amount of text and a sufficiently long text.
[0181] Furthermore, the target model provided in this application embodiment is a pre-trained model. To adapt to downstream tasks, the target model (structure and parameters) can be fine-tuned according to the needs of different downstream tasks. The following describes this process in conjunction with several downstream tasks:
[0182] Let the downstream task be a document question-and-answer task, such as Figure 8 As shown ( Figure 8 (This is another structural diagram of the target model provided in the embodiments of this application). The decoder in the target model is still connected to the coordinate converter and the text converter. Figure 8 The structure of the target model shown and Figure 7The target model shown can have the same structure. In this case, the BOS input to the decoder is the prompt (question) entered by the user. The decoder will output the location information of the answer in vector form and the answer in vector form based on the features of the image entered by the user. After being transformed by a coordinate converter, the coordinates of the answer can be obtained, and after being transformed by a text converter, the text form of the answer can be obtained. Then, the coordinates of the answer and the text form of the answer can be attached to the image entered by the user and returned to the user.
[0183] For example, such as Figure 9 As shown ( Figure 9 (This is a schematic diagram of a document Q&A provided in an embodiment of this application). The user inputs an image of a high-speed rail ticket and asks, "Where did X Haipeng go?" The target model can process the image and obtain the text answer "Zhalanping Station" and the coordinates of the answer "Zhalanping Station" in the image. Then, "Zhalanping Station" can be selected in the image and displayed with corresponding highlighted text for the user to view.
[0184] Let the downstream task be an information extraction task, such as Figure 10 As shown ( Figure 10 (This is another structural schematic diagram of the target model provided in the embodiments of this application), and the decoder in the target model is connected to the information extractor (that is, Figure 7 In the target model shown, the converter is replaced by an information extractor. At this point, the BOS input to the decoder is still a meaningless vector representation. The decoder, based on the features of the user-input image, outputs the location information of the target information in vector form, as well as the target information in vector form. After comprehensive processing by the information extractor, the target information in text form can be obtained.
[0185] For example, such as Figure 11 As shown ( Figure 11 (This is a schematic diagram illustrating information extraction provided in an embodiment of this application.) The user inputs an image of a high-speed rail ticket. The target model can process the image to obtain information such as time, destination, name, departure point, and train number in text form, and then return this information to the user for browsing.
[0186] Furthermore, the target model provided in the embodiments of this application can be compared with a model provided by a related technology, and the comparison results are shown in Table 1:
[0187] Table 1
[0188] Model Decoding output Long document decoding length computational complexity Related technologies text Unrestricted higher Examples of this application Text + Location Information Unrestricted Lower
[0189] As shown in Table 1, the target model provided in this application embodiment can not only decode more types of output, but also has lower computational complexity, effectively reducing the power consumption of the decoder in the model and speeding up the decoding speed.
[0190] Furthermore, the target model provided in the embodiments of this application (e.g., the ours in Table 2) can be compared with models provided by other related technologies (e.g., other models besides the ours in Table 2, such as BERT, etc.), and the comparison results are shown in Table 2:
[0191] Table 2
[0192]
[0193] As shown in Table 2, the target model provided in this application embodiment has better performance.
[0194] In this embodiment, when it is necessary to extract target text from a target image, a target image containing multiple texts can first be obtained and input into a target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process these features to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within it. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from the multiple texts presented in the target image in this manner is usually the correct text.
[0195] Furthermore, the final output of the target model not only includes the target text in character form (text), but also the coordinates of the area occupied by the target text in the target image. These two types of information can be superimposed on the target image and returned to the user for browsing. In this way, the target text required by the user can be provided through a visual interactive interface, and the basis for the target model to extract the target text can also be explained to the user.
[0196] Furthermore, the output length of the target model is unlimited. In this way, even if the user needs to obtain a long text or a large amount of text, the target model can meet the user's needs, thereby improving the user experience.
[0197] The above is a detailed description of the text acquisition method provided in the embodiments of this application. The model training method provided in the embodiments of this application will be introduced below. Figure 12 A schematic flowchart of the model training method provided in the embodiments of this application is shown below. Figure 12 As shown, the method includes:
[0198] 1201. Obtain the target image, which contains multiple texts.
[0199] In this embodiment, when the model to be trained needs to be trained, a batch of training data can be obtained first. This batch of training data includes target images, wherein the content presented by the target images includes multiple texts. For the multiple texts contained in the target images, all the real characters of the target texts are known, and the real coordinates of the target texts in the area occupied by the target texts in the target images are also known.
[0200] 1202. The target image is processed by the model to be trained to obtain the location information of the target text in the target image and the target text. Multiple texts contain the target text. The model to be trained is used to: encode the target image to obtain the features of the target image; obtain the location information based on the features; and obtain the target text based on the features and the location information.
[0201] After obtaining the target image, it can be input into the model to be trained. The model can then encode the target image to obtain its features. Next, based on these features, the model can extract the location information of the target text within the target image. Finally, based on both the target image features and the location information of the target text, the model can extract the target text.
[0202] In one possible implementation, the model to be trained is used to decode the first to the ith vector representation of the positional information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the positional information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the positional information is obtained by decoding the preset vector representation based on features.
[0203] In one possible implementation, the model to be trained is used to decode the first to the j-th vector representation of the target text based on the feature-based location information to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on the feature.
[0204] In one possible implementation, the target text includes a first text and a second text, and the location information includes the first location information of the first text in the target image and the second location information of the second text in the target image. The model to be trained is used to: decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, and the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, and the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0205] In one possible implementation, the model to be trained is used to: decode the first positional information, the first vector representation of the first text to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the positional information based on features; and decode the first positional information, the first text, the second positional information, the first vector representation of the second text to the t-th vector representation of the second text based on features, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first positional information, the first text, and the second positional information based on features.
[0206] In one possible implementation, the model to be trained is also used to transform all vector representations of the location information to obtain the (predicted) coordinates of the region occupied by the target text in the target image.
[0207] In one possible implementation, the model to be trained is also used to transform all vector representations of the target text to obtain all (predicted) characters of the target text.
[0208] In one possible implementation, the transformation performed by the model to be trained on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0209] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0210] For an explanation of step 1202, please refer to [link / reference]. Figure 5 The relevant descriptions of steps 502 to 504 in the illustrated embodiment will not be repeated here.
[0211] 1203. Based on the target text, train the model to be trained to obtain the target model.
[0212] Once the target text is obtained, the model to be trained can be trained based on the target text, thereby obtaining... Figure 5 The target model in the illustrated embodiment.
[0213] Specifically, the target model can be trained in the following ways:
[0214] After obtaining all the characters of the target text and the coordinates of the region occupied by the target text in the target image, a first loss function can be used to calculate the first loss for both the characters and the actual characters of the target text. The first loss can be obtained using the following formula:
[0215]
[0216] In the above formula, L Read Let X be the first loss, and y be the feature of the target image. r-1 For the (r-1)th character of the target text, y r Let y' be the r-th character of the target text. r Let be the r-th real character of the target text. Therefore, the first loss can be used to indicate the difference between the characters in the target text and the real characters in the target text.
[0217] Next, a second loss can be obtained by calculating the coordinates of the region occupied by the target text in the target image and the true coordinates of the region occupied by the target text in the target image using a preset second loss function. The second loss can be obtained using the following formula:
[0218]
[0219] In the above formula, L Locate This is the second loss. Let be the (u-1)th coordinate of the region occupied by the s-th target text in the target image. Let u be the coordinate of the region occupied by the s-th target text in the target image. Let be the u-th true coordinate of the region occupied by the s-th target text in the target image. Therefore, the second loss can be used to indicate the difference between the coordinates of the region occupied by the target text in the target image and the true coordinates of the region occupied by the target text in the target image.
[0220] Then, the first loss and the second loss can be superimposed to obtain the target loss. The parameters of the model to be trained can then be updated using the target loss to obtain the updated model to be trained. The updated model can then be trained using the next batch of training data until the model training conditions are met (e.g., the target loss converges, etc.), thus obtaining the target model.
[0221] The target text trained in this embodiment of the application has a text acquisition function. Specifically, when it is necessary to extract target text from a target image, a target image containing multiple texts can be acquired first, and the target image can be input into the target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process the features of the target image to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text within the target image to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within the target image. This comprehensive consideration allows for a full and accurate understanding of the content of the target image. Therefore, the target text extracted by the target model from the multiple texts presented in the target image in this manner is usually the correct text.
[0222] The above is a detailed description of the text acquisition method and model training method provided in the embodiments of this application. The following will introduce the text acquisition device and model training device provided in the embodiments of this application. Figure 13 A schematic diagram of the structure of the text acquisition device provided in the embodiments of this application is shown below. Figure 13 As shown, the device includes:
[0223] The first acquisition module 1301 is used to acquire a target image, which contains multiple texts;
[0224] Encoding module 1302 is used to encode the target image to obtain the features of the target image;
[0225] The second acquisition module 1303 is used to acquire the location information of the target text in the target image based on features, wherein multiple texts contain the target text;
[0226] The third acquisition module 1304 is used to acquire target text based on features and location information.
[0227] In this embodiment, when it is necessary to extract target text from a target image, a target image containing multiple texts can first be obtained and input into a target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process these features to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within it. This comprehensive consideration allows for a thorough and accurate understanding of the target image's content. Therefore, the target text extracted by the target model from the multiple texts presented in the target image in this manner is usually the correct text.
[0228] In one possible implementation, the second acquisition module 1303 is used to decode the first vector representation to the ith vector representation of the position information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the position information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the position information is obtained by decoding the preset vector representation based on features.
[0229] In one possible implementation, the third acquisition module 1304 is used to decode the location information, from the first vector representation to the j-th vector representation of the target text, based on features, to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on features.
[0230] In one possible implementation, the target text includes first text and second text, and the location information includes first location information of the first text in the target image and second location information of the second text in the target image. The second acquisition module 1303 is used to decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, where the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and to decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, where the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0231] In one possible implementation, the third acquisition module 1304 is used to decode the first position information and the first vector representation of the first text up to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the position information based on features; based on features, the third acquisition module 1304 decodes the first position information, the first text, the second position information, and the first vector representation of the second text up to the t-th vector representation of the second text, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first position information, the first text, and the second position information based on features.
[0232] In one possible implementation, the device further includes: a first transformation module for transforming all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
[0233] In one possible implementation, the device further includes a second conversion module for converting all vector representations of the target text to obtain all characters of the target text.
[0234] In one possible implementation, the transformation performed by the target model on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0235] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0236] Figure 14 A schematic diagram of the model training apparatus provided in the embodiments of this application is shown below. Figure 14 As shown, the device includes:
[0237] The acquisition module 1401 is used to acquire a target image, which contains multiple texts;
[0238] The processing module 1402 is used to process the target image through the model to be trained to obtain the location information of the target text in the target image and the target text. Multiple texts contain the target text. The model to be trained is used to: encode the target image to obtain the features of the target image; obtain the location information based on the features; and obtain the target text based on the features and the location information.
[0239] Training module 1403 is used to train the model to be trained based on the target text to obtain the target model.
[0240] The target text trained in this embodiment of the application has a text acquisition function. Specifically, when it is necessary to extract target text from a target image, a target image containing multiple texts can be acquired first, and the target image can be input into the target model. Next, the target model can encode the target image to obtain its features. Then, the target model can process the features of the target image to obtain the positional information of the target text within the target image. Finally, the target model can further process the features of the target image and the positional information of the target text within the target image to obtain the target text. Thus, the target text is successfully extracted from the target image. In the aforementioned process, when understanding the content of the target image, the target model considers not only the features of the target image but also the positional information of the target text within the target image. This comprehensive consideration allows for a full and accurate understanding of the content of the target image. Therefore, the target text extracted by the target model from the multiple texts presented in the target image in this manner is usually the correct text.
[0241] In one possible implementation, the model to be trained is used to decode the first to the ith vector representation of the positional information of the target text in the target image based on features, to obtain the (i+1)th vector representation of the positional information, i = 1, ..., X-1, X ≥ 1. The first vector representation of the positional information is obtained by decoding the preset vector representation based on features.
[0242] In one possible implementation, the model to be trained is used to decode the first to the j-th vector representation of the target text based on the feature-based location information to obtain the (j+1)-th vector representation of the target text, where j = 1, ..., Y-1, Y ≥ 1. The first vector representation of the target text is obtained by decoding the location information based on the feature.
[0243] In one possible implementation, the target text includes a first text and a second text, and the location information includes the first location information of the first text in the target image and the second location information of the second text in the target image. The model to be trained is used to: decode the first vector representation to the i-th vector representation of the first location information based on features to obtain the (i+1)-th vector representation of the first location information, i = 1, ..., X-1, X ≥ 1, and the first vector representation of the first location information is obtained by decoding a preset vector representation based on features; and decode the first location information, the first text, and the first vector representation to the k-th vector representation of the second location information based on features to obtain the (k+1)-th vector representation of the first location information, k = 1, ..., Z-1, Z ≥ 1, and the first vector representation of the second location information is obtained by decoding the first location information and the first text based on features.
[0244] In one possible implementation, the model to be trained is used to: decode the first positional information, the first vector representation of the first text to the j-th vector representation of the first text based on features, to obtain the (j+1)-th vector representation of the first text, j = 1, ..., Y-1, Y ≥ 1, and the first vector representation of the first text is obtained by decoding the positional information based on features; and decode the first positional information, the first text, the second positional information, the first vector representation of the second text to the t-th vector representation of the second text based on features, to obtain the (t+1)-th vector representation of the second text, t = 1, ..., U-1, U ≥ 1, and the first vector representation of the second text is obtained by decoding the first positional information, the first text, and the second positional information based on features.
[0245] In one possible implementation, the model to be trained is also used to transform all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
[0246] In one possible implementation, the model to be trained is also used to transform all vector representations of the target text to obtain all characters of the target text. Training module 1403 is used to train the model to be trained based on the characters and coordinates to obtain the target model.
[0247] In one possible implementation, the transformation performed by the model to be trained on the target text and location information can be at least one of the following: feature extraction based on recurrent neural networks, feature extraction based on multilayer perceptrons, and feature extraction based on temporal convolutional networks.
[0248] In one possible implementation, the coordinates of the region are at least one of the following: the coordinates of the top-left vertex and the bottom-right vertex of the region; or, the coordinates of the top-right vertex and the bottom-left vertex of the region; or, the coordinates of the four corner vertices of the region; or, the coordinates of the top-left vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, and the center point of the region; or, the coordinates of the top-right vertex, the top-left vertex, and the center point of the region; or, the coordinates of the bottom-right vertex, the bottom-left vertex, and the center point of the region; or, the coordinates of the top-right vertex, the bottom-right vertex, the top-left vertex, the bottom-left vertex, and the center point of the region.
[0249] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of this application, and the resulting technical effects are the same as those of the method embodiment of this application. For details, please refer to the description in the method embodiment shown above in the embodiment of this application, and it will not be repeated here.
[0250] This application also relates to an execution device. Figure 15 This is a schematic diagram of the execution device provided in an embodiment of this application. Figure 15 As shown, the execution device 1500 can specifically be a mobile phone, tablet, laptop, smart wearable device, server, etc., and is not limited here. Among them, the execution device 1500 can be deployed with... Figure 13 The text acquisition device described in the corresponding embodiment is used to implement Figure 5 This corresponds to the text acquisition function in the embodiment. Specifically, the execution device 1500 includes: a receiver 1501, a transmitter 1502, a processor 1503, and a memory 1504 (wherein the execution device 1500 may have one or more processors 1503). Figure 15 (Taking a processor as an example), processor 1503 may include application processor 15031 and communication processor 15032. In some embodiments of this application, receiver 1501, transmitter 1502, processor 1503 and memory 1504 may be connected via a bus or other means.
[0251] Memory 1504 may include read-only memory and random access memory, and provides instructions and data to processor 1503. A portion of memory 1504 may also include non-volatile random access memory (NVRAM). Memory 1504 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0252] Processor 1503 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.
[0253] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 1503. Processor 1503 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 1503 or by instructions in software form. Processor 1503 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 1503 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1504. Processor 1503 reads the information in memory 1504 and, in conjunction with its hardware, completes the steps of the above method.
[0254] Receiver 1501 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 1502 can be used to output digital or character information through the first interface; transmitter 1502 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 1502 may also include a display device such as a display screen.
[0255] In one embodiment of this application, the processor 1503 is used to... Figure 5 The target model in the corresponding embodiment obtains the target text from the target image.
[0256] This application also relates to a training device. Figure 16 This is a schematic diagram of the structure of a training device provided in an embodiment of this application. Figure 16 As shown, the training device 1600 is implemented by one or more servers. The training device 1600 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1616 (e.g., one or more processors) and memory 1632, and one or more storage media 1630 (e.g., one or more mass storage devices) for storing application programs 1642 or data 1644. The memory 1632 and storage media 1630 can be temporary or persistent storage. The program stored in the storage media 1630 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the training device. Furthermore, the CPU 1616 may be configured to communicate with the storage media 1630 and execute the series of instruction operations in the storage media 1630 on the training device 1600.
[0257] The training device 1600 may also include one or more power supplies 1626, one or more wired or wireless network interfaces 1650, one or more input / output interfaces 1658; or, one or more operating systems 1641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0258] Specifically, the training equipment can perform Figure 12 The model training method in the corresponding embodiment is used to obtain the target model.
[0259] This application also relates to a computer storage medium storing a program for signal processing, which, when run on a computer, causes the computer to perform steps as performed by the aforementioned execution device, or causes the computer to perform steps as performed by the aforementioned training device.
[0260] This application also relates to a computer program product that stores instructions that, when executed by a computer, cause the computer to perform steps as performed by the aforementioned execution device, or to perform steps as performed by the aforementioned training device.
[0261] The execution device, training device, or terminal device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip within the execution device to execute the data processing method described in the above embodiments, or to cause the chip within the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0262] For details, please refer to Figure 17 , Figure 17 This is a schematic diagram of the chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 1700. The NPU 1700 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 1703, which is controlled by the controller 1704 to retrieve matrix data from the memory and perform multiplication operations.
[0263] In some implementations, the arithmetic circuit 1703 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 1703 is a two-dimensional pulsating array. The arithmetic circuit 1703 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1703 is a general-purpose matrix processor.
[0264] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1702 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1701 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 1708.
[0265] Unified memory 1706 is used to store input and output data. Weight data is directly transferred to weight memory 1702 via Direct Memory Access Controller (DMAC) 1705. Input data is also transferred to unified memory 1706 via DMAC.
[0266] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1709.
[0267] The Bus Interface Unit (BIU) 1717 is used by the instruction fetch memory 1709 to fetch instructions from external memory, and also by the memory access controller 1705 to fetch the original data of the input matrix A or the weight matrix B from external memory.
[0268] The DMAC is mainly used to move input data from external memory DDR to unified memory 1706, or to weight data to weight memory 1702, or to input data to input memory 1701.
[0269] The vector computation unit 1707 includes multiple processing units that further process the output of the computation circuit 1703 when needed, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as Batch Normalization, pixel-level summation, and upsampling of the predicted label plane.
[0270] In some implementations, the vector computation unit 1707 can store the processed output vector in the unified memory 1706. For example, the vector computation unit 1707 can apply a linear function, or a nonlinear function, to the output of the computation circuit 1703, such as linearly interpolating the predicted label plane extracted from the convolutional layer, or, for example, accumulating a vector of values to generate activation values. In some implementations, the vector computation unit 1707 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1703, for example, for use in subsequent layers of the neural network.
[0271] The instruction fetch buffer 1709 connected to the controller 1704 is used to store the instructions used by the controller 1704.
[0272] Unified memory 1706, input memory 1701, weighted memory 1702, and instruction fetch memory 1709 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.
[0273] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of the above program.
[0274] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0275] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0276] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0277] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A text acquisition method, characterized in that, The method is implemented through a target model, and the method includes: Acquire a target image, which contains multiple texts; The target image is encoded to obtain its features; Based on the aforementioned features, the location information of the target text in the target image is obtained, wherein the plurality of texts contains the target text; Based on the features and the location information, the target text is obtained; The target text includes first text and second text, and the location information includes first location information of the first text in the target image and second location information of the second text in the target image. Obtaining the location information of the target text in the target image based on the feature includes: Based on the aforementioned features, the first vector representation to the ith vector representation of the first location information is decoded to obtain the (i+1)th vector representation of the first location information, where i=1,...,X-1, X≥1. The first vector representation of the first location information is obtained by decoding a preset vector representation based on the aforementioned features. Based on the aforementioned features, the first location information, the first text, and the first to the kth vector representations of the second location information are decoded to obtain the (k+1)th vector representation of the first location information, where k=1,...,Z-1, Z≥1. The first vector representation of the second location information is obtained by decoding the first location information and the first text based on the aforementioned features.
2. The method according to claim 1, characterized in that, The step of obtaining the location information of the target text in the target image based on the features includes: Based on the aforementioned features, the first vector representation to the ith vector representation of the location information of the target text in the target image are decoded to obtain the (i+1)th vector representation of the location information, where i=1,...,X-1, X≥1. The first vector representation of the location information is obtained by decoding a preset vector representation based on the aforementioned features.
3. The method according to claim 2, characterized in that, The process of obtaining the target text based on the features and the location information includes: Based on the features, the location information and the first vector representation to the j-th vector representation of the target text are decoded to obtain the (j+1)-th vector representation of the target text, where j=1,...,Y-1, Y≥1. The first vector representation of the target text is obtained by decoding the location information based on the features.
4. The method according to claim 1, characterized in that, The process of obtaining the target text based on the features and the location information includes: Based on the features, the first location information and the first vector representation of the first text to the j-th vector representation of the first text are decoded to obtain the (j+1)-th vector representation of the first text, j=1,...,Y-1, Y≥1. The first vector representation of the first text is obtained by decoding the location information based on the features. Based on the aforementioned features, the first location information, the first text, the second location information, and the first vector representation to the t-th vector representation of the second text are decoded to obtain the (t+1)-th vector representation of the second text, where t=1,...,U-1, U≥1. The first vector representation of the second text is obtained by decoding the first location information, the first text, and the second location information based on the aforementioned features.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The coordinates of the region occupied by the target text in the target image are obtained by transforming all vector representations of the location information.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Transform all vector representations of the target text to obtain all characters of the target text.
7. A model training method, characterized in that, The method includes: Acquire a target image, which contains multiple texts; The target image is processed by a model to be trained to obtain the location information of the target text in the target image and the target text, wherein the plurality of texts contain the target text. The model to be trained is used to: encode the target image to obtain the features of the target image; obtain the location information based on the features; and obtain the target text based on the features and the location information. Based on the target text, the model to be trained is trained to obtain the target model; The target text includes first text and second text, the location information includes first location information of the first text in the target image and second location information of the second text in the target image, and the model to be trained is used for: Based on the aforementioned features, the first vector representation to the ith vector representation of the first location information is decoded to obtain the (i+1)th vector representation of the first location information, where i=1,...,X-1, X≥1. The first vector representation of the first location information is obtained by decoding a preset vector representation based on the aforementioned features. Based on the aforementioned features, the first location information, the first text, and the first to the kth vector representations of the second location information are decoded to obtain the (k+1)th vector representation of the first location information, where k=1,...,Z-1, Z≥1. The first vector representation of the second location information is obtained by decoding the first location information and the first text based on the aforementioned features.
8. The method according to claim 7, characterized in that, The model to be trained is used to decode the first vector representation to the ith vector representation of the position information of the target text in the target image based on the features, to obtain the (i+1)th vector representation of the position information, i=1,...,X-1, X≥1. The first vector representation of the position information is obtained by decoding a preset vector representation based on the features.
9. The method according to claim 8, characterized in that, The model to be trained is used to decode the location information, the first vector representation to the j-th vector representation of the target text, based on the features, to obtain the (j+1)-th vector representation of the target text, where j=1,...,Y-1, Y≥1. The first vector representation of the target text is obtained by decoding the location information based on the features.
10. The method according to claim 7, characterized in that, The model to be trained is used for: Based on the features, the first location information and the first vector representation of the first text to the j-th vector representation of the first text are decoded to obtain the (j+1)-th vector representation of the first text, j=1,...,Y-1, Y≥1. The first vector representation of the first text is obtained by decoding the location information based on the features. Based on the aforementioned features, the first location information, the first text, the second location information, and the first vector representation to the t-th vector representation of the second text are decoded to obtain the (t+1)-th vector representation of the second text, where t=1,...,U-1, U≥1. The first vector representation of the second text is obtained by decoding the first location information, the first text, and the second location information based on the aforementioned features.
11. The method according to any one of claims 7 to 10, characterized in that, The model to be trained is also used to transform all vector representations of the location information to obtain the coordinates of the region occupied by the target text in the target image.
12. The method according to claim 11, characterized in that, The model to be trained is also used to transform all vector representations of the target text to obtain all characters of the target text; The step of training the model to be trained based on the target text to obtain the target model includes: Based on the characters and coordinates, the model to be trained is trained to obtain the target model.
13. A text acquisition device, characterized in that, The device includes a target model, and the device comprises: The first acquisition module is used to acquire a target image, wherein the target image contains multiple texts; An encoding module is used to encode the target image to obtain the features of the target image; The second acquisition module is used to acquire the location information of the target text in the target image based on the features, wherein the plurality of texts includes the target text; The third acquisition module is used to acquire the target text based on the features and the location information; The target text includes first text and second text, and the location information includes first location information of the first text in the target image and second location information of the second text in the target image. The second acquisition module is used for: Based on the aforementioned features, the first vector representation to the ith vector representation of the first location information is decoded to obtain the (i+1)th vector representation of the first location information, where i=1,...,X-1, X≥1. The first vector representation of the first location information is obtained by decoding a preset vector representation based on the aforementioned features. Based on the aforementioned features, the first location information, the first text, and the first to the kth vector representations of the second location information are decoded to obtain the (k+1)th vector representation of the first location information, where k=1,...,Z-1, Z≥1. The first vector representation of the second location information is obtained by decoding the first location information and the first text based on the aforementioned features.
14. A model training device, characterized in that, The device includes: The acquisition module is used to acquire a target image, wherein the target image contains multiple texts; The processing module is used to process the target image using a model to be trained, to obtain the location information of the target text in the target image and the target text, wherein the plurality of texts include the target text, and the model to be trained is used to: encode the target image to obtain the features of the target image; obtain the location information based on the features; and obtain the target text based on the features and the location information. The training module is used to train the model to be trained based on the target text to obtain the target model; The target text includes first text and second text, the location information includes first location information of the first text in the target image and second location information of the second text in the target image, and the model to be trained is used for: Based on the aforementioned features, the first vector representation to the ith vector representation of the first location information is decoded to obtain the (i+1)th vector representation of the first location information, where i=1,...,X-1, X≥1. The first vector representation of the first location information is obtained by decoding a preset vector representation based on the aforementioned features. Based on the aforementioned features, the first location information, the first text, and the first to the kth vector representations of the second location information are decoded to obtain the (k+1)th vector representation of the first location information, where k=1,...,Z-1, Z≥1. The first vector representation of the second location information is obtained by decoding the first location information and the first text based on the aforementioned features.
15. A text acquisition device, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, wherein when the code is executed, the text acquisition device performs the method as described in any one of claims 1 to 12.
16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 12.
17. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 12.
Citation Information
Patent Citations
Text recognition method and device, medium and electronic equipment
CN114758342A