Table Image Matching Method, System, Medium and Device
By extracting text features, title features and visual features from table pictures, and calculating cosine similarity using the cross attention mechanism, the problem of low matching accuracy in table pictures in the prior art is solved, and higher matching accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510255580.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the prior art, the table image matching method mainly relies on visual feature extraction, resulting in too low matching accuracy, especially when there are blurred, distorted, overlapped or irregularly formatted text, lines and cells in the table, the matching accuracy decreases.
The matching features are extracted by extracting text features, title features and visual features from the two table pictures to be compared, and inputting these features into the cross attention mechanism. Then, the cosine similarity of the matching features is calculated to determine the degree of matching of the table picture.
It improves the accuracy and accuracy of table image matching, enhances the robustness of picture matching technology, and thus improves matching efficiency and security.
Smart Images

Figure CN119760166B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of natural language processing, and relates to a matching method, in particular to a table image matching method, system, medium and device. Background Art
[0002] In actual application scenarios, it is often necessary to compare the matching degree of two table images. Table image matching mainly refers to a subtask of image matching, and image matching is actually image retrieval. That is: when a user inputs an image and wants to find images with similar content, through image matching technology, images with similar content are found and returned to the user; for a more specific application, it is image search by image.
[0003] With the development of deep learning technology, especially the application of convolutional neural networks (CNNs) in the field of image recognition, the matching technology of table images has been significantly improved. Existing algorithms can relatively accurately identify the text, lines and cell structures in the table, and then realize the automatic extraction and matching of the table. This enables the matching of table images to be applied in many scenarios, such as extracting table data from scanned documents or photos, and comparing and merging information between different tables.
[0004] Currently, one of the core problems of table image matching methods is that the matching accuracy is not high enough. Although deep learning algorithms have made great progress in image recognition, in the face of complex table structures and variable image quality, it is often difficult for the algorithms to accurately extract and match the key information in the table. Especially when there are blurred, distorted, overlapping or non-standard formatted text, lines and cells in the table, the matching accuracy will further decrease, resulting in incorrect data extraction and matching results.
[0005] Therefore, considering the above reasons, especially in the prior art, the matching of table images is mainly completed by extracting visual features of the table images. However, this method has the problem of low matching accuracy.
[0006] Therefore, how to improve the accuracy of table image matching is an urgent problem to be solved at present. Summary of the Invention
[0007] In view of the above-mentioned disadvantages of the prior art, the purpose of this application is to provide a table image matching method, system, medium and device, which are used to solve the problem that in the existing image matching technology, the accuracy of image matching is too low due to only using visual features for matching.
[0008] To achieve the above and other related objectives, in a first aspect, the present application provides a method for matching table pictures, including the following steps: obtaining a first table picture and a second table picture to be compared; extracting features based on the first table picture to obtain first matching features; extracting features based on the second table picture to obtain second matching features; calculating the pairwise similarity between the first matching features and the second matching features; and determining the matching degree between the first table picture and the second table picture based on the pairwise similarity.
[0009] In an implementation manner of the first aspect, extracting features based on the first table picture to obtain first matching features includes the following steps: extracting features from the first table picture to obtain first title features, first visual features, and first text features respectively; using the first title features as K vectors, the first visual features as V vectors, and the first text features as Q vectors, and obtaining first matching features based on the cross-attention mechanism.
[0010] In an implementation manner of the first aspect, the first visual features of the first table picture are extracted based on the VIT model.
[0011] In an implementation manner of the first aspect, extracting features from the first table picture to obtain first text features includes the following steps: performing optical character recognition on the first table picture to obtain text line positions and text line contents; extracting all the text, vertical text, and horizontal text in the first table picture according to the text line positions; extracting text semantic features of all the text, the vertical text, and the horizontal text based on the BERT model to obtain first all-text semantic features, first vertical-text semantic features, and first horizontal-text semantic features; and calculating the average value of the first all-text semantic features, the first vertical-text semantic features, and the first horizontal-text semantic features to obtain first text features.
[0012] In an implementation manner of the first aspect, extracting features from the first table picture to obtain first title features includes the following steps: combining text prompts, inputting the vertical text and horizontal text in the text line positions into a large language model to obtain column titles and row titles; extracting semantic features of the column titles and the row titles based on the text encoder of CLIP to obtain column title semantic features and row title semantic features; and calculating the average value of the column title semantic features and the row title semantic features to obtain first title features.
[0013] In an implementation of the first aspect, taking the first title feature as the K vector, the first visual feature as the V vector, and the first text feature as the Q vector, obtaining the first matching feature based on the cross-attention mechanism includes the following steps: taking the first title feature as the K vector, the first visual feature as the V vector, and the first text feature as the Q vector, obtaining the first fusion feature based on the cross-attention mechanism; inputting the first fusion feature into a multi-layer perceptron to obtain a first fusion feature vector; inputting the first fusion feature vector into the sigmoid function to obtain the first matching feature.
[0014] In an implementation of the first aspect, feature extraction based on the second table picture to obtain the second matching feature includes the following steps: performing feature extraction on the second table picture to respectively obtain a second title feature, a second visual feature, and a second text feature; taking the second title feature as the K vector, the second visual feature as the V vector, and the second text feature as the Q vector, and obtaining the second matching feature based on the cross-attention mechanism.
[0015] In a second aspect, the present application provides a table picture matching system, including: an acquisition module, a first feature extraction module, a second feature extraction module, and a picture matching module; the acquisition module is configured to acquire a first table picture and a second table picture to be compared; the first feature extraction module is configured to perform feature extraction based on the first table picture to obtain a first matching feature; the second feature extraction module is configured to perform feature extraction based on the second table picture to obtain a second matching feature; the picture matching module is configured to calculate the pairwise similarity between the first matching feature and the second matching feature, and determine the matching degree between the first table picture and the second table picture based on the pairwise similarity.
[0016] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by an electronic device, it implements the table picture matching method according to any one of the first aspects of the present application.
[0017] In the last aspect, the present application provides a table picture matching device, including: a processor and a memory. The memory is used to store a computer program; the processor is connected to the memory and is configured to execute the computer program stored in the memory, so that the table picture matching device executes the table picture matching method.
[0018] As described above, the table picture matching method, system, medium, and device of the present invention have the following beneficial effects:
[0019] The present application provides a method for matching table images. Feature extraction is respectively performed on two different table images to be compared to obtain the text features, title features, and visual features of each table image. Then, the text features, title features, and visual features corresponding to the table images are input into a cross-attention mechanism to extract the matching features corresponding to different table images respectively. Next, the cosine similarity of the matching features of the two table images to be compared is calculated. Then, based on the cosine similarity, the matching degree of the two table images is determined, which can effectively improve the accuracy and precision of the matching of different table images. At the same time, the matching method adopted in the present application can improve the robustness of different image matching technologies, and further improve the matching efficiency and security. Description of the Drawings
[0020] Figure 1 It shows a schematic diagram of a scenario in an embodiment of the electronic device of the present application.
[0021] Figure 2 It shows a schematic flowchart of the table image matching method described in the present application in an embodiment.
[0022] Figure 3 It shows a schematic flowchart of the matching process of different table images in an embodiment of the table image matching method described in the present application.
[0023] Figure 4 It shows a schematic flowchart of S2 in the table image matching method described in the present application.
[0024] Figure 5 It shows a schematic flowchart of the first text feature extraction in an embodiment of the table image matching method described in the present application.
[0025] Figure 6 It shows a schematic flowchart of the first title feature extraction in an embodiment of the table image matching method described in the present application.
[0026] Figure 7 It shows a schematic flowchart of S22 in the table image matching method described in the present application.
[0027] Figure 8 It shows a schematic flowchart of S3 in the table image matching method described in the present application.
[0028] Figure 9 It shows a schematic diagram of the principle structure of the table image matching system described in the present application in an embodiment.
[0029] Figure 10 It shows a schematic diagram of the principle structure of the table image matching device described in the present application in an embodiment.
[0030] Description of Component Labels
[0031] 11 Mobile phone 12 Tablet computer 13 Laptop computer 91 Acquisition module 92 First feature extraction module 93 Second feature extraction module 94 Image matching module 101 Processor 102 Memory S1~S4 Step Detailed Implementation Manner
[0032] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0033] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0034] The following embodiments of the present application provide a method, system, medium, and device for table image matching. Feature extraction is performed separately from two different table images to be compared to obtain the text features, title features, and visual features of each table image. Then, the text features, title features, and visual features corresponding to the table images are input into the cross-attention mechanism to extract the matching features corresponding to different table images; then, the cosine similarity of the matching features of the two table images to be compared is calculated; then, based on the cosine similarity, the matching degree of the two table images is determined, which can effectively improve the matching accuracy of different table images.
[0035] The table image matching method of the present application can be applied to an electronic device as shown in Figure 1 The electronic device described in the present application may include a mobile phone 11 with a wireless charging function, a tablet computer 12, a laptop computer 13, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The specific type of the electronic device is not limited in the embodiments of the present application.
[0036] For example, the electronic device may be a station (STAION, ST) in a WLAN with wireless charging function, a cellular phone with wireless charging function, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless charging function, a computing device or other processing device, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system, as well as next-generation communication systems, such as mobile terminals in a 5G network, mobile terminals in a future evolved Public Land Mobile Network (PLMN), or mobile terminals in a future evolved Non-terrestrial Network (NTN), etc.
[0037] For example, the electronic device may communicate with the network and other devices through wireless communication. The above wireless communication may use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0038] The following will combine the accompanying drawings in the embodiments of the present application to describe in detail the table picture matching method provided in the embodiments of the present application.
[0039] Please refer to Figure 2 and Figure 3 , which are respectively the schematic flowchart of the table picture matching method described in the present application in an embodiment and the schematic flowchart of the matching process of different table pictures in an embodiment of the table picture matching method described in the present application. As Figure 2 and Figure 3 shown, this embodiment provides a table picture matching method.
[0040] The table picture matching method specifically includes the following steps:
[0041] S1, obtain a first table picture and a second table picture to be compared.
[0042] In this embodiment, the table pictures in the present application specifically refer to pictures that only contain tables (tables of course include text). Table pictures can be obtained in various ways such as reading from local files, downloading from the Internet, or uploading by users.
[0043] It should be noted that the first table picture and the second table picture can be read through image processing tools (such as: OpenCV, PIL, etc.).
[0044] S2, perform feature extraction based on the first table picture to obtain a first matching feature. Please refer to Figure 4 , which shows the schematic flowchart of S2 in the table picture matching method described in the present application. As Figure 4 shown, S2 includes the following steps:
[0045] S21, perform feature extraction on the first table picture to obtain a first title feature, a first visual feature, and a first text feature respectively.
[0046] In this embodiment, the first visual feature of the first table picture is extracted based on the VIT model.
[0047] Specifically, first, preprocess the first table picture, such as: adjusting the size of the table picture, scaling, denoising, binarizing, normalizing, etc., to ensure that the quality and format of the input data meet the requirements of the VIT model. Such processing helps to improve the performance and stability of the model. Then, divide the preprocessed table picture into several small pieces, that is: each small piece is a patch. Next, for each image patch, convert it into an embedding vector in a high-dimensional space through linear projection (such as: fully connected layer) to capture the visual features of the patch. Then use the self-attention mechanism in the Transformer model to process these embedding vectors. By calculating the correlation between each patch, the model can learn the dependency relationships between different patches and extract visual features that contribute to the final task accordingly. Finally, aggregate the feature vectors of all patches to form a feature representation of the complete table picture, that is, the first visual feature of the first table picture.
[0048] Similarly, extract features from the first table picture to obtain the first text feature. Please refer to Figure 5 , which shows the schematic flowchart of the first text feature extraction in an embodiment of the table picture matching method described in this application.
[0049] In this embodiment, perform character recognition on the first table picture to obtain the text line position and text line content; according to the text line position, extract all the characters, vertical characters, and horizontal characters in the first table picture; based on the BERT model, extract the text semantic features of all the characters, the vertical characters, and the horizontal characters to obtain the first all-character semantic feature, the first vertical-character semantic feature, and the first horizontal-character semantic feature; calculate the average value of the first all-character semantic feature, the first vertical-character semantic feature, and the first horizontal-character semantic feature to obtain the first text feature.
[0050] Specifically, the OCR (Optical Character Recognition) method is used to recognize the input first table picture, and the text line positions and text line contents can be obtained. Then, according to the text line positions, the text lines of each column are found and named vertical texts (there are multiple columns); the text lines of each row are found and named horizontal texts (there are multiple rows). Then, for all texts (arranged in the order from top to bottom and from left to right of the text lines, and each text line is spliced with a comma), vertical texts and horizontal texts, the BERT model is used to extract text semantic features, and three types of text semantic features can be obtained respectively, namely: the first all-text semantic feature, the first vertical-text semantic feature, and the first horizontal-text semantic feature. Next, the average values of the first all-text semantic feature, the first vertical-text semantic feature, and the first horizontal-text semantic feature are calculated, and the obtained results are used as the first text feature.
[0051] The BERT (Bidirectional Encoder Representation from Transformers) model is a pre-trained language model based on Transformer. It is pre-trained with a large amount of unlabeled corpus and then fine-tuned with a small amount of labeled corpus to complete specific NLP tasks. The core of the BERT model is based on the Transformer structure, especially its Encoder part, which is composed of a stack of multiple Transformer encoders (BertLayer), and each layer contains a self-attention mechanism and a feed-forward neural network. The pre-training tasks of BERT include Masked Language Model (MLM) and Next Sentence Prediction (NSP). The former randomly masks some words in the sentence and predicts the masked content, and the latter predicts whether two sentences come from the same document. These pre-training tasks enable BERT to learn sentence-level semantic information and thus perform well in various NLP tasks.
[0052] Feature extraction is performed on the first table picture to obtain a first title feature. Please refer to Figure 6 , which shows a schematic flowchart of the first title feature extraction in an embodiment of the table picture matching method described in this application.
[0053] In this embodiment, in combination with the text prompt, the vertical text and horizontal text extracted from the text line positions are input into a large language model to obtain column headers and row headers; the semantic features of the column headers and the row headers are extracted by the text encoder of CLIP to obtain column header semantic features and row header semantic features; the average value of the column header semantic features and the row header semantic features is calculated to obtain the first header feature.
[0054] Specifically, according to the horizontal text and vertical text obtained in the above steps, as the input of the large language model LLM, the input is set as: "Find the row headers in the text. The text content is like the horizontal text. If there is a header, directly answer the header name without explanation. If not, answer none without explanation."; after inputting this text into the LLM, the answer (row header) is obtained; similarly, the column header can be obtained. Then, the semantic features of the column header and the row header are respectively extracted by the CLIP text encoder to obtain column header semantic features and row header semantic features, and then the average value of the column header semantic features and the row header semantic features is calculated, and the result is the first header feature.
[0055] For example: after inputting the horizontal text into the LLM model, the answer is obtained, that is, the row header; after inputting the vertical text into the LLM model, the answer is obtained, that is, the column header. Preprocess the extracted row header and column header content, such as removing noise, punctuation marks, etc., and perform necessary text cleaning and normalization processing. This helps to improve the understanding and representation ability of the text encoder for the text content. Then, the preprocessed column header and row header are input into the CLIP text encoder to extract key features and convert them into vectors to obtain column header semantic features and row header semantic features. Then, the average value of the extracted column header semantic features and row header semantic features is calculated to obtain the first header feature.
[0056] It should be noted that the CLIP text encoder is trained based on a large-scale dataset. Therefore, when extracting text features, it can capture more extensive and in-depth semantic information, which makes the text feature matrix extracted based on the CLIP text encoder have better performance in tasks such as text classification, clustering, and retrieval.
[0057] S22, take the first header feature as the K vector, take the first visual feature as the V vector, take the first text feature as the Q vector, and obtain the first matching feature based on the cross-attention mechanism. Please refer to Figure 7 , which shows the flow schematic diagram of S22 in the table picture matching method described in this application.
[0058] In this embodiment, the first title feature is used as the K vector, the first visual feature is used as the V vector, and the first text feature is used as the Q vector, and a first fusion feature is obtained based on the cross-attention mechanism; the first fusion feature is input into a multi-layer perceptron to obtain a first fusion feature vector; the first fusion feature vector is input into a sigmoid function to obtain the first matching feature.
[0059] Specifically, first, the first title feature is used as the K vector, the first visual feature is used as the V vector, and the first text feature is used as the Q vector and input into the MLP. The similarity between the first title feature and the first text feature is calculated through the attention layer, that is: the dot product, cosine similarity or other similarity measurement methods are used to calculate the similarity between the first title feature and the first text feature; then, these similarity values are normalized into attention weights through the softmax function. Then, the first visual feature is weighted according to these similarities to obtain the weighted feature, that is: the first fusion feature.
[0060] Further, the first fusion feature is used as the input and fed into the pre-constructed MLP model. The MLP will perform forward propagation, and through the weights and activation functions between layers, the input features are transformed and processed. After calculation by the MLP model, one or more feature vectors will be output. These output vectors are the results of the original fusion feature mapped by the MLP, and usually represent the abstract representation of the input features at a higher level. The output of the MLP is used as the input of the sigmoid function. The sigmoid function will map the input value to the range between 0 and 1 to obtain a representation of probability or activation degree. The output of the sigmoid function is the first matching feature. This feature combines the original fusion feature and the non-linear transformation ability of the MLP, as well as the mapping characteristics of the sigmoid function.
[0061] It should be noted that the specific implementation steps may be adjusted according to the specific tasks and datasets. For example, it may be necessary to optimize the parameters of the MLP to achieve better performance; or it may be necessary to add other processing steps before or after the sigmoid function, such as: normalization, dimensionality reduction, etc.
[0062] S3. Feature extraction is performed based on the second table picture to obtain a second matching feature. Please refer to Figure 8 , which shows the flow schematic diagram of S3 in the table picture matching method described in this application. As Figure 8 shown, S3 includes the following steps:
[0063] S31. Feature extraction is performed on the second table picture to respectively obtain a second title feature, a second visual feature, and a second text feature.
[0064] In this embodiment, the second visual feature of the second table picture is extracted based on the VIT model.
[0065] Specifically, first, preprocess the second table picture, such as: adjusting the size of the table picture, scaling, denoising, binarizing, normalizing, etc., to ensure that the quality and format of the input data meet the requirements of the VIT model. Such processing helps to improve the performance and stability of the model. Then, divide the preprocessed table picture into several small pieces, that is: each small piece is a patch. Next, for each image patch, convert it into an embedding vector in a high-dimensional space through linear projection (such as: fully connected layer) to capture the visual features of the patch. Then use the self-attention mechanism in the Transformer model to process these embedding vectors. By calculating the correlation between each patch, the model can learn the dependency relationships between different patches and extract the visual features that contribute to the final task accordingly. Finally, aggregate the feature vectors of all patches to form a feature representation of the complete table picture, that is, the second visual feature of the second table picture.
[0066] Similarly, extract the second text feature from the second table picture.
[0067] In this embodiment, perform optical character recognition on the second table picture to obtain the text line positions and text line contents; according to the text line positions, extract all the characters, vertical characters, and horizontal characters in the second table picture; extract the text semantic features of all the characters, the vertical characters, and the horizontal characters based on the BERT model to obtain the second all-character semantic feature, the second vertical-character semantic feature, and the second horizontal-character semantic feature; calculate the average values of the second all-character semantic feature, the second vertical-character semantic feature, and the second horizontal-character semantic feature to obtain the second text feature.
[0068] Specifically, use the OCR method to recognize the input second table picture to obtain the text line positions and text line contents. Then, according to the text line positions, find the text lines in each column, named vertical characters (there are multiple columns); find the text lines in each row, named horizontal characters (there are multiple rows). Then, for all the characters (arranged in the order of text lines from top to bottom and from left to right, with commas spliced between each text line), vertical characters, and horizontal characters, use the BERT model to extract text semantic features, and 3 types of text semantic features can be obtained respectively, that is: the second all-character semantic feature, the second vertical-character semantic feature, and the second horizontal-character semantic feature. Next, calculate the average values of the second all-character semantic feature, the second vertical-character semantic feature, and the second horizontal-character semantic feature, and the result obtained is used as the second text feature.
[0069] Extract features from the second table picture to obtain second title features.
[0070] In this embodiment, in combination with text prompts, the vertical and horizontal text extracted from the text line positions are input into a large language model to obtain column titles and row titles; the semantic features of the column titles and the row titles are extracted based on the text encoder of CLIP to obtain column title semantic features and row title semantic features; the average value of the column title semantic features and the row title semantic features is calculated to obtain second title features.
[0071] Specifically, according to the horizontal and vertical text obtained in the above steps, as the input of the large language model LLM, the input is set as: "Find the row titles in the text. The text content is like the horizontal text. If there is a title, directly answer the title name without explanation. If not, answer none without explanation."; After inputting this text into the LLM, the answer (row title) is obtained; Similarly, the column title can be obtained. Then, the semantic features of the column title and the row title are respectively extracted through the CLIP text encoder to obtain column title semantic features and row title semantic features, and then the average value calculation is performed on the column title semantic features and the row title semantic features. The result obtained is the second title feature.
[0072] S32. Use the second title feature as the K vector, the second visual feature as the V vector, and the second text feature as the Q vector to obtain a second matching feature based on the cross-attention mechanism.
[0073] In this embodiment, use the second title feature as the K vector, the second visual feature as the V vector, and the second text feature as the Q vector to obtain a second fusion feature based on the cross-attention mechanism; input the second fusion feature into a multi-layer perceptron to obtain a second fusion feature vector; input the second fusion feature vector into the sigmoid function to obtain a second matching feature.
[0074] Specifically, first, input the second title feature as the K vector, the second visual feature as the V vector, and the second text feature as the Q vector into the MLP. Calculate the similarity between the second title feature and the second text feature through the attention layer, that is: use dot product, cosine similarity or other similarity measurement methods to calculate the similarity between the second title feature and the second text feature; then, normalize these similarity values into attention weights through the softmax function. Then, weight the second visual feature according to these similarities to obtain the weighted feature, that is: the second fusion feature.
[0075] Further, the second fusion feature is used as the input and fed into the pre-constructed MLP model. The MLP performs forward propagation, transforming and processing the input features through the weights and activation functions between layers. After calculation, the MLP model outputs a feature vector. The output of the MLP is used as the input to the sigmoid function. The sigmoid function maps the input value to the range between 0 and 1, obtaining a representation of probability or activation degree. The output of the sigmoid function is the second matching feature. This feature combines the original fusion feature, the non-linear transformation ability of the MLP, and the mapping characteristics of the sigmoid function.
[0076] It should be noted that the entire process from the table picture to the extraction of the matching feature corresponding to the table picture can be regarded as a feature extraction process. The feature extraction methods and the entire process for obtaining the matching features of the first table picture and the second table picture are exactly the same. Therefore, the feature extraction processes of the first table picture and the second table picture are consistent.
[0077] S4. Calculate the pairwise similarity between the first matching feature and the second matching feature, and determine the matching degree between the first table picture and the second table picture based on the pairwise similarity.
[0078] In this embodiment, as can be seen from the foregoing steps, by performing feature extraction on the first table picture, the first matching feature is obtained; by performing feature extraction on the second table picture, the second matching feature is obtained. Then, calculate the cosine similarity between the first matching feature and the second matching feature; determine the matching degree between the first table picture and the second table picture through the value of the cosine similarity; for example, the closer the value of the cosine similarity is to 1, the more matching it is; on the contrary, it indicates a greater difference between the first table picture and the second table picture.
[0079] Therefore, after completing all the above steps and judging the matching degree of the two table pictures through the finally calculated cosine similarity value, the matching result of the two table pictures can be obtained.
[0080] The table image matching method provided by this application extracts features from two different table images to be compared, obtains the text features, title features, and visual features of each table image, and then inputs the text features, title features, and visual features corresponding to the table images into the cross-attention mechanism to extract the matching features corresponding to different table images; then calculates the cosine similarity of the matching features of the two table images to be compared; then determines the matching degree of the two table images based on the cosine similarity, which can effectively improve the accuracy and precision of the matching of different table images. At the same time, the matching method adopted by this application can improve the robustness of different image matching technologies, and further improve the matching efficiency and security.
[0081] The protection scope of the table image matching method described in the embodiments of this application is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding or subtracting steps of the prior art and replacing steps according to the principle of this application is included in the protection scope of this application.
[0082] This embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the Figure 1 table image matching method described above.
[0083] At any possible combination level of technical details, this application can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions are loaded to enable the processor to implement various aspects of this application.
[0084] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding device, such as punched cards or raised structures in grooves storing instructions thereon, and any suitable combination of the above. The computer-readable storage medium used here is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0085] The computer-readable programs described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device. The computer program instructions for performing the operations of the present application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user computer, partially on the user computer, executed as a stand-alone software package, partially on the user computer and partially on the computer, or entirely on the computer or server. In the case of a computer, the computer may be connected to the user computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet). In some embodiments, by using the status information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present application.
[0086] An embodiment of the present application also provides a table picture matching system. The table picture matching system can implement the table picture matching method described in the present application. However, the implementation devices of the table picture matching method described in the present application include, but are not limited to, the structures of the table picture matching system listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of the present application are included in the protection scope of the present application.
[0087] The table picture matching system provided in this embodiment will be described in detail below with reference to the drawings.
[0088] This embodiment provides a table picture matching system, including:
[0089] Please refer to Figure 9 , which shows a schematic diagram of the principle structure of the table picture matching system described in the present application in an embodiment. AsFigure 9 As shown in Figure 9 , the table image matching system includes: an acquisition module 91, a first feature extraction module 92, a second feature extraction module 93, and an image matching module 94.
[0090] The acquisition module 91 is configured to acquire a first table image and a second table image to be compared.
[0091] In this embodiment, the table image in this application specifically refers to an image that only contains a table (the table naturally includes text). The table image can be acquired in various ways such as reading from a local file, downloading from the Internet, or uploading by the user.
[0092] The first feature extraction module 92 is configured to extract features based on the first table image to obtain first matching features.
[0093] Feature extraction is performed on the first table image to obtain a first title feature, a first visual feature, and a first text feature respectively.
[0094] In this embodiment, the first visual feature of the first table image is extracted based on the VIT model.
[0095] Specifically, first, preprocess the first table image, such as: adjusting the size of the table image, scaling, denoising, binarizing, performing normalization operations, etc., to ensure that the quality and format of the input data meet the requirements of the VIT model. Such processing helps to improve the performance and stability of the model. Then, divide the preprocessed table image into several small pieces, that is: each small piece is a patch. Next, for each image patch, convert it into an embedding vector in a high-dimensional space through linear projection (such as: a fully connected layer) to capture the visual features of the patch. Then use the self-attention mechanism in the Transformer model to process these embedding vectors. By calculating the correlation between each patch, the model can learn the dependency relationships between different patches and extract visual features that contribute to the final task accordingly. Finally, converge the feature vectors of all patches to form a feature representation of the complete table image, that is, the first visual feature of the first table image.
[0096] Similarly, feature extraction is performed on the first table image to obtain a first text feature.
[0097] In this embodiment, optical character recognition is performed on the first table picture to obtain the text line positions and text line contents; according to the text line positions, all the characters, vertical characters, and horizontal characters in the first table picture are extracted; based on the BERT model, the text semantic features of all the characters, the vertical characters, and the horizontal characters are extracted to obtain the first all-character semantic features, the first vertical-character semantic features, and the first horizontal-character semantic features; the average values of the first all-character semantic features, the first vertical-character semantic features, and the first horizontal-character semantic features are calculated to obtain the first text features.
[0098] Feature extraction is performed on the first table picture to obtain the first title features.
[0099] In this embodiment, in combination with text prompts, the vertical characters and the horizontal characters are input into a large language model to obtain column headers and row headers; based on the text encoder of CLIP, the semantic features of the column headers and the row headers are extracted to obtain column header semantic features and row header semantic features; the average values of the column header semantic features and the row header semantic features are calculated to obtain the first title features.
[0100] Taking the first title features as the K vectors, taking the first visual features as the V vectors, and taking the first text features as the Q vectors, the first matching features are obtained based on the cross-attention mechanism.
[0101] In this embodiment, taking the first title features as the K vectors, taking the first visual features as the V vectors, and taking the first text features as the Q vectors, the first fusion features are obtained based on the cross-attention mechanism; the first fusion features are input into a multi-layer perceptron to obtain the first fusion feature vectors; the first fusion feature vectors are input into the sigmoid function to obtain the first matching features.
[0102] It should be noted that the above entire implementation steps can be regarded as a whole, that is, the feature extraction process of the first table picture, or this process can be regarded as a large feature extraction module; thus, using the same large feature extraction module, feature extraction is performed on the second table picture to obtain the second matching features, details are as follows.
[0103] The second feature extraction module 93 is used to perform feature extraction based on the second table picture to obtain the second matching features.
[0104] Feature extraction is performed on the second table picture to obtain the second title features, the second visual features, and the second text features respectively.
[0105] In this embodiment, the second visual features of the second table picture are extracted based on the VIT model.
[0106] Similarly, feature extraction is performed on the second table picture to obtain second text features.
[0107] In this embodiment, character recognition is performed on the second table picture to obtain the text line positions and text line contents; according to the text line positions, all characters, vertical characters, and horizontal characters in the second table picture are extracted; based on the BERT model, the text semantic features of all the characters, the vertical characters, and the horizontal characters are extracted to obtain second all-character semantic features, second vertical-character semantic features, and second horizontal-character semantic features; the average values of the second all-character semantic features, the second vertical-character semantic features, and the second horizontal-character semantic features are calculated to obtain second text features.
[0108] Feature extraction is performed on the second table picture to obtain second title features.
[0109] In this embodiment, in combination with text prompts, the vertical characters and the horizontal characters are input into a large language model to obtain column titles and row titles; based on the text encoder of CLIP, the semantic features of the column titles and the row titles are extracted to obtain column title semantic features and row title semantic features; the average value of the column title semantic features and the row title semantic features is calculated to obtain second title features.
[0110] Furthermore, the second title feature is used as the K vector, the second visual feature is used as the V vector, and the second text feature is used as the Q vector, and a second matching feature is obtained based on the cross-attention mechanism.
[0111] In this embodiment, the second title feature is used as the K vector, the second visual feature is used as the V vector, and the second text feature is used as the Q vector, and a second fusion feature is obtained based on the cross-attention mechanism; the second fusion feature is input into a multi-layer perceptron to obtain a second fusion feature vector; the second fusion feature vector is input into a sigmoid function to obtain a second matching feature.
[0112] The picture matching module 94 is used to calculate the pairwise similarities between the first matching feature and the second matching feature, and determine the matching degree between the first table picture and the second table picture based on the pairwise similarities.
[0113] In this embodiment, as can be seen from the foregoing steps, by performing feature extraction on the first table picture, a first matching feature is obtained; by performing feature extraction on the second table picture, a second matching feature is obtained. Then, the cosine similarity between the first matching feature and the second matching feature is calculated; based on the value of the cosine similarity, the matching degree between the first table picture and the second table picture is determined; for example, the closer the value of the cosine similarity is to 1, the more matching; on the contrary, it means that the gap between the first table picture and the second table picture is greater.
[0114] Therefore, after completing all the above steps, the cosine similarity value obtained through the final calculation is used to judge the matching degree of the two table pictures, and the matching result of the two table pictures can be obtained.
[0115] Building a table picture matching model for a table picture matching system can effectively improve the matching accuracy of different table pictures.
[0116] It should be noted that it should be understood that the division of each module of the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the x module can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, it can also be stored in the memory of the above system in the form of program code, and called and executed by a certain processing element of the above system to perform the functions of the above x module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0117] The above modules can be one or more integrated circuits configured to implement the above method, for example: one or more application specific integrated circuits (ASICs), or, one or more microprocessors (digital signal processors, DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a program code scheduled by a processing element, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0118] Please refer to Figure 10 , which shows the schematic diagram of the principle structure of the table picture matching device described in this application in an embodiment. As Figure 10As shown in the figure, this embodiment provides a table picture matching device, which includes: a processor 101 and a memory 102; the memory 102 is used to store a computer program; the processor 101 is connected to the memory 102 and is used to execute the computer program stored in the memory 102, so that the table picture matching device executes each step of the table picture matching method described above.
[0119] Preferably, the memory may include a random access memory (Random Access Memory, RAM for short), and may also include a non-volatile memory, such as at least one disk memory.
[0120] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, CPU for short), a network processor (Network Processor, NP for short), etc.; it may also be a digital signal processor (Digital Signal Processing, DSP for short), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC for short), a field programmable gate array (Field Programmable Gate Array, FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0121] In summary, the table picture matching method, system, medium and device provided by this application have the following beneficial effects:
[0122] The table picture matching method provided by this application extracts features from two different table pictures to be compared, obtains the text features, title features and visual features of each table picture, and then inputs the text features, title features and visual features corresponding to the table pictures into the cross-attention mechanism to extract the matching features corresponding to different table pictures; then calculates the cosine similarity of the matching features of the two table pictures to be compared; then, based on the cosine similarity, determines the matching degree of the two table pictures, which can effectively improve the accuracy and accuracy of the matching of different table pictures. At the same time, the matching method adopted by this application can improve the robustness of different picture matching technologies, and further improve the matching efficiency and security.
[0123] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in the present application should still be covered by the claims of the present application.
Claims
1. A table image matching method, characterized in that: The following steps are involved: Obtain a first table image and a second table image to be compared; Extracting features based on the first table image to obtain a first matching feature; including: extracting features from the first table image to obtain a first title feature, a first visual feature, and a first text feature; using the first title feature as a K vector, the first visual feature as a V vector, and the first text feature as a Q vector, and obtaining the first matching feature based on a cross-attention mechanism; Extracting features based on the second table image to obtain a second matching feature; including: extracting features from the second table image to obtain a second title feature, a second visual feature, and a second text feature respectively; using the second title feature as a K vector, the second visual feature as a V vector, and the second text feature as a Q vector, and obtaining a second matching feature based on a cross-attention mechanism; Calculating the pairwise similarity between the first matching feature and the second matching feature, and determining the matching degree between the first table image and the second table image based on the pairwise similarity; The extraction of the text features includes the following steps: performing text recognition on the table image to obtain the text row position and the text row content; extracting all the text, vertical text, and horizontal text in the table image according to the text row position; extracting the text semantic features of all the text, the vertical text, and the horizontal text based on the BERT model to obtain the semantic features of all the text, the semantic features of the vertical text, and the semantic features of the horizontal text; calculating the average values of the semantic features of all the text, the semantic features of the vertical text, and the semantic features of the horizontal text to obtain the text features; The extraction of the title feature includes: extracting features of the table image based on the column title and the row title to obtain the title feature.
2. The table image matching method according to claim 1, characterized in that: The first visual feature of the first table image is extracted based on the VIT model.
3. The table image matching method according to claim 1, characterized in that: Extracting features from the first table image to obtain a first title feature includes the following steps: Combined with the text prompt, the vertical text and the horizontal text in the text line position are input into the large language model to obtain the column title and the row title; A text encoder based on CLIP extracts semantic features of the column header and the row header to obtain column header semantic features and row header semantic features; An average value of the column title semantic feature and the row title semantic feature is calculated to obtain a first title feature.
4. The table image matching method according to claim 1, characterized in that: Taking the first title feature as a K vector, taking the first visual feature as a V vector, taking the first text feature as a Q vector, and obtaining a first matching feature based on a cross attention mechanism includes the following steps: Taking the first title feature as a K vector, taking the first visual feature as a V vector, taking the first text feature as a Q vector, and obtaining a first fusion feature based on a cross attention mechanism; Inputting the first fused feature into a multi-layer perceptron to obtain a first fused feature vector; The first fused feature vector is input into a sigmoid function to obtain the first matching feature.
5. A table image matching system, characterized in that: include: An acquisition module, a first feature extraction module, a second feature extraction module and a picture matching module; The acquisition module is used to acquire the first table image and the second table image to be compared; The first feature extraction module is used to extract features based on the first table image to obtain a first matching feature; including: extracting features from the first table image to obtain a first title feature, a first visual feature, and a first text feature; using the first title feature as a K vector, the first visual feature as a V vector, and the first text feature as a Q vector, and obtaining the first matching feature based on a cross attention mechanism; The second feature extraction module is used to extract features based on the second table image to obtain a second matching feature; including: extracting features from the second table image to obtain a second title feature, a second visual feature, and a second text feature respectively; using the second title feature as a K vector, the second visual feature as a V vector, and the second text feature as a Q vector, and obtaining the second matching feature based on a cross attention mechanism; The image matching module is used to calculate the pairwise similarity between the first matching feature and the second matching feature, and determine the matching degree between the first table image and the second table image based on the pairwise similarity; The extraction of the text features includes the following steps: performing text recognition on the table image to obtain the text row position and the text row content; extracting all the text, vertical text, and horizontal text in the table image according to the text row position; extracting the text semantic features of all the text, the vertical text, and the horizontal text based on the BERT model to obtain the semantic features of all the text, the semantic features of the vertical text, and the semantic features of the horizontal text; calculating the average values of the semantic features of all the text, the semantic features of the vertical text, and the semantic features of the horizontal text to obtain the text features; The extraction of the title feature includes: extracting features of the table image based on the column title and the row title to obtain the title feature.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by an electronic device, the table image matching method described in any one of claims 1 to 4 is implemented.
7. A table image matching device, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is connected to the memory and is used to execute the computer program stored in the memory so that the table image matching device executes the table image matching method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-source table automatic matching method and device based on semantics and storage medium
CN117252183A
Table identification method, identification model training method and related products
CN119399784A