A scene text retrieval model, method and computer device based on text detection and semantic matching
By introducing a scene text search method based on text detection and semantic matching in the cross-modal search model, and using the scene text information in the image for searching, the problems of search accuracy and inefficiency in the prior art are solved, and a more efficient and accurate cross-modal search effect is achieved.
Patent Information
- Application Number
- CN202210716996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-06-23
AI Technical Summary
Existing cross-modal search methods fail to effectively use the scene text information in the image for retrieval, resulting in inaccuracy and inefficiency of search.
Design a scene text retrieval model based on text detection and semantic matching. By extracting scene text features and image description text features in the image, using a multi-layer perceptron and stacked cross attention mechanism to perform similarity calculations, so as to achieve semantic alignment between images and text.
The search accuracy of the cross-modal retrieval model on the data containing scene text is improved, the influence of noise information is reduced, and the accuracy and efficiency of the model are enhanced.
Smart Images

Figure CN115017266B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cross-modal retrieval, and in particular relates to a scene text retrieval model, method and computer device based on text detection and semantic matching. Background Art
[0002] With the rapid growth of multimodal data on the Internet, many challenging tasks have been paid attention to and studied in order to effectively analyze and utilize data of different modalities. Cross-modal retrieval is one of these tasks and a hot research topic at home and abroad. Given a picture or image text description, the corresponding image text description or picture is retrieved.
[0003] Images are important data carriers and the amount of data is huge. In daily life, text information can be quickly searched through keyword search. However, it is difficult to search for information carried in images, which greatly reduces the efficiency of obtaining information from images. Cross-modal retrieval can solve this problem to a certain extent. At the same time, images may also contain some important text information. If this part of information can also be obtained through retrieval, the accuracy of cross-modal retrieval can be greatly improved. Thanks to the development of image recognition, optical character recognition (OCR) technology has gradually matured and has been widely used in character detection and recognition.
[0004] The main research direction of the present invention is the study of cross-modal retrieval algorithms between images and texts. Although a large number of feasible algorithms have emerged in this field in recent years and the accuracy has been continuously improved, the existing methods do not use the text information in the image to retrieve the corresponding text description, or use the text description to search for pictures with text information. Therefore, cross-modal retrieval still faces huge challenges in OCR-based retrieval. At present, there is still a unified model that uses scene text in images to perform cross-modal retrieval tasks. The research purpose of the present invention is to improve the efficiency and accuracy of information retrieval by using scene text in pictures to perform retrieval tasks on the basis of existing cross-modal methods. Therefore, the present invention designs a scene text retrieval method based on text detection and semantic matching to improve the retrieval accuracy of the cross-modal retrieval model on data containing scene text. Summary of the invention
[0005] The purpose of the present invention is to utilize scene text information in images on the basis of cross-modal retrieval, learn the features of scene text, and explore the interaction between deep-level picture semantics and image description features. This can better align the image and text semantically, thereby improving the accuracy of retrieval. The present invention further uses a deep learning method to extract the features of scene text and image description text on the basis of the existing cross-modal retrieval model, by using an image text extraction system and natural semantic analysis tools, etc., to perform similarity calculation, thereby obtaining a cross-modal retrieval model that can be used for scene text retrieval. The technical solution for implementing the present invention is as follows:
[0006] A scene text cross-modal retrieval model based on text detection and semantic matching is obtained through the following steps:
[0007] S1, extracts the regional features of the image and the word-level features of the image description, and maps the two features to a common semantic space through a multi-layer perceptron.
[0008] S2, uses cosine similarity to calculate the similarity between the two, optimizes the model through triple loss function, and finally obtains cross-modal retrieval similarity.
[0009] S3, uses the Rosetta OCR image text extraction system to extract the text information in the image, namely the scene text, and uses fastText to extract the word features of the scene text.
[0010] S4, uses StanfordCoreNlp to process the image description, selects words that meet the semantic requirements and extracts word features through fastText.
[0011] S5, considering the different levels of text features, uses word and sentence level features to calculate similarity, and uses the stacked cross attention mechanism to make the model model the semantic relationship between scene text and image description, and the three similarities are weighted to obtain the final similarity between scene text and image description.
[0012] S6, integrates the model of step S2 and the model of S5 into a unified framework to obtain a scene text retrieval model based on text detection and semantic matching.
[0013] Step S1-1, use the pre-trained FasterRCNN to extract the regional features of the image, and input the extracted features into the multi-layer perceptron, map them to the common feature space, and obtain the feature
[0014]
[0015] Step S1-2: put the input image description into Bi-GRU to extract the features of each word, and input the extracted features into the multi-layer perceptron, map them to the common feature space, and obtain the features
[0016]
[0017] Step S2-1, using cosine similarity to measure the similarity of the image and image description features obtained in steps S1-1 and S1-2, and finally obtaining a similarity result of cross-modal retrieval.
[0018] Step S3-1, extracting scene text from the image through Rosetta OCR;
[0019] Step S3-2, preprocessing the extracted scene text to filter out words that are too short, symbols and other noise.
[0020] In step S3-3, the scene text obtained in step S3-2 is screened by part-of-speech analysis, and a 300-dimensional feature vector is extracted from the final scene text through a pre-trained fastText model.
[0021] Step S4-1, performing natural semantic analysis on the image description text to obtain the part of speech of each word;
[0022] Step S4-2, by performing part-of-speech analysis on the scene text and comparing it with the image description text, the parts of speech that are effective for model retrieval are screened out, and the corresponding words are input into fastText;
[0023] In step S4-3, fastText is pre-trained on the training set and the input words are extracted into 300-dimensional feature vectors.
[0024] Step S5-1, the input picture and image description text are used to obtain the word features of the scene text and the image description text according to step S3 and step S4. Two sentence-level semantic features are obtained by averaging the word features, and the sentence-level similarity between the scene text and the image description text is calculated using cosine similarity.
[0025] Step S5-2, the input picture and image description text are used to obtain the word features of the scene text and the image description text according to step S2 and step S3. A k*l matrix is constructed based on the two texts, where k and l represent the sentence length of the scene text and the sentence length of the image description text respectively. The word-level similarity between the scene text and the image description text is obtained through matrix operation.
[0026] Step S5-3, the input picture and image description text are used to obtain the word features of the scene text and image description text according to step S2 and step S3. The attention similarity between the two texts is calculated by using the stacked cross attention mechanism. The three similarities are weighted to obtain the final similarity between the scene text and the image description.
[0027] Step S6-1, integrate the model of step S2 and the model of S5 into a unified framework to obtain a scene text retrieval model based on text detection and semantic matching.
[0028] S=S c +λ 1 S sum
[0029] Among them, λ 1 is the balancing parameter, S c is the cross-modal retrieval similarity obtained in step S2, S sum It is the sentence level, word level and stacked cross attention similarity obtained by step S5.
[0030] The method of using the above model for cross-modal retrieval is as follows:
[0031] First, load the trained neural network parameters into the model. For image description text retrieval using images, the image to be retrieved and the image description text in the dataset are passed into the model as input. The model calculates the similarity between the image and other image description texts and outputs the top ten image description texts that are most similar to the input image.
[0032] For image retrieval using image description text, the image description text to be retrieved and the images in the dataset are passed into the model. The model calculates the similarity between the image description text and other images and outputs the top ten images most similar to the input image description text.
[0033] The above-mentioned scene text cross-modal retrieval model and method based on text detection and semantic matching can be placed in a computer device, either in the form of executing instruction program code or in the form of storing program code.
[0034] Beneficial effects of the present invention:
[0035] (1) The scene text retrieval method based on text detection and semantic matching proposed in the present invention enables the cross-modal retrieval model to utilize the text information in the image by learning the mapping relationship between scene text and image description text.
[0036] (2) By preprocessing the text data, the noise information in the data is removed, and by using fastText to extract text features, the impact of scene text recognition errors is reduced, thereby improving the accuracy of the model.
[0037] (3) The present invention uses multiple levels of text similarity to perform retrieval, so that the model can be trained at both the overall and local levels, so that the model can effectively improve the accuracy of retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flowchart of scene text retrieval based on text detection and semantic matching. DETAILED DESCRIPTION
[0039] This method first extracts image and text features to build a basic cross-modal retrieval module; uses the Rosetta OCR image text extraction system to extract text information in the image, namely scene text, and uses fastText to extract text features in the image; uses StanfordCoreNlp to process the scene text, selects words that meet semantic requirements, and extracts text features through fastText; considering the different levels of text features, similarity calculations are performed using word and sentence level features, and the attention mechanism is used for calculation, so that the model models the semantic relationship between scene text and image description; the obtained cross-modal retrieval similarity and the similarity between scene text and image description are integrated to obtain a scene text retrieval model based on text detection and semantic matching. Through the above steps, a multi-level and unified scene text retrieval method is obtained. By introducing scene text, the retrieval accuracy on images containing text can be effectively improved on the existing cross-modal retrieval model.
[0040] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0041] like Figure 1 As shown, the specific implementation of the present invention includes the following steps:
[0042] Step S1, extracting the regional features of the image and the word-level features of the image description, and mapping the two features to a common semantic space through a multi-layer perceptron;
[0043] The step S1 further comprises the following steps:
[0044] Step S1-1, given an image I, use the pre-trained FasterRCNN to detect n regions of interest r in the image i , and extract the corresponding regional features f i Then use a multi-layer perceptron to transform the image region features f iMapped to the common feature space to get v i :
[0045] v i =MLP v (f i )
[0046] Among them, MLP v Represents the multi-layer perceptron corresponding to the image, and the obtained image features are expressed as Where n represents the number of regions of interest detected using FasterRCNN.
[0047] Step S1-2, given a sentence T (image description text), for the i-th word in the sentence, use the one-hot encoding w i Indicates the position of the word in the vocabulary, using the mapping matrix W e w i Mapped into a 300-dimensional vector, represented as x i =W e w u ,i∈[1,m], where m represents the number of words in the sentence. Use Bi-GRU to transform x u The Bi-GRU consists of a forward GRU that maps w 1 To w m Read sentence T, the specific implementation is as follows:
[0048]
[0049] and a backward GRU, from w m To w 1 Read sentence T, the specific implementation is as follows:
[0050]
[0051] Final word feature e i By and By averaging the fusion, the word features are fused w i The context information of the surrounding sentences is expressed as Then use a multi-layer perceptron to map the image description to a common feature space to obtain e i :
[0052] e i =MLP e (g i )
[0053] Among them, MLP e Represents the multi-layer perceptron corresponding to the image, and the text feature of the input image is represented as
[0054] Step S2, using cosine similarity to calculate the similarity between the two, optimizing the model through triple loss function, and finally obtaining the cross-modal retrieval similarity;
[0055] The step S2 further comprises the following steps:
[0056] Step S2-1, using cosine similarity to calculate the similarity of the image feature V and text feature E obtained in steps S1-1 and S1-2, and obtain the similarity result of cross-modal retrieval. visual (·) and text aggregator f text (·) is aggregated, and the image feature V and text feature E are further embedded to obtain the aggregated features α, β:
[0057]
[0058] The similarity between image I and image description text T is calculated by cosine similarity, denoted by S c Represents the cross-modal retrieval similarity, and obtains the first sub-model, S c It is expressed as:
[0059]
[0060] Train the first sub-model using triplet loss:
[0061]
[0062] Where Δ is a hyperparameter and (v,e) represents the dataset The positive sample pairs in represents the hardest negative sample of v, represents the hardest negative sample of t, [x] + ≡max(0,x), the ternary sorting loss is used to shorten the distance between positive sample pairs, and v′ and e′ are intermediate variables.
[0063] Step S3, using the Rosetta OCR image text extraction system to extract text information in the image, namely, scene text, and using fastText to extract word features of the scene text;
[0064] The step S3 further comprises the following steps:
[0065] Step S3-1: for a given image I, input the image into the Rosetta OCR image text extraction system to extract all OCR tokens in the image, where OCR tokens are words recognized from the image. For each input image, use OCR to extract word text, i.e., scene text.
[0066] Step S3-2, preprocessing the extracted scene text, first cleaning the data, deleting symbols, single characters and other recognized text.
[0067] Step S3-3, perform part-of-speech analysis on the scene text obtained in step S3-2, and send the scene text to StanfordCoreNlp for semantic analysis. The image descriptions and Rosetta OCR scene texts in the image descriptions are analyzed. The image descriptions and scene texts have a large number of words with the same part of speech, which correspond to the part-of-speech tags defined in StanfordCoreNlp, namely NN (noun, common, singular or large number), NNS (noun, common, plural), NNP (noun, proper, singular), CD (number, cardinality), and JJ (adjective or numeral, ordinal). Therefore, the scene texts preprocessed by S3-2 are screened for part of speech, and words with word parts of speech included in the above five parts of speech are selected, and finally k scene texts are selected for subsequent tasks.
[0068] The final scene text is extracted into a 300-dimensional feature vector through the pre-trained fastText model. The word feature representation of the scene text obtained by fastText is:
[0069] Step S4, using StanfordCoreNlp to process the image description, select words that meet the semantic requirements and extract word features through fastText;
[0070] The step S4 further comprises the following steps:
[0071] Step S4-1, performing part-of-speech analysis on the image description text, ie, the unprocessed word text, and sending the image description text to StanfordCoreNlp for semantic analysis to obtain the part-of-speech tag corresponding to each word.
[0072] Step S4-2, similar to step S3-3, is to The image descriptions and Rosetta OCR scene texts in the image description text are used for data analysis, and the image description words with the parts of speech of NN, NNS, NNP, CD, and JJ are selected from the image description text.
[0073] Step S4-3, extract a 300-dimensional feature vector from the final image description text through the pre-trained fastText model. The word feature representation of the input image description text obtained by fastText is
[0074] S5, considering the different levels of text features, uses word and sentence level features to calculate similarity, and uses the stacked cross attention mechanism to make the model model the semantic relationship between scene text and image description, and obtains the final similarity between scene text and image description by weighting the three similarities;
[0075] The step S5 further comprises the following steps:
[0076] Step S5-1: The input image and image description text are used to obtain the word features O and P of the scene text and the image description text according to step S3 and step S4. The words of the scene text are averaged and fused to obtain the sentence level representation O of the current scene text. s ,
[0077]
[0078] The words of the image description text are averaged and fused to obtain the sentence level representation P of the current image description text. s ,
[0079]
[0080] The cosine similarity is used to calculate the sentence-level similarity between the scene text and the image description text, and S s Represents the sentence-level similarity, expressed as:
[0081]
[0082] Step S5-2, the input picture and image description text, according to step S3 and step S4, obtain the word features O and P of the scene text and image description text, and form the word-level features of the scene text and image description text by splicing the word features, which are expressed as [o i ,...,o k ],[p i ,...,p l ]. The two features are matrices of size k*300 and l*300 respectively. Construct a k*l cosine similarity matrix S based on the two text features. t , each value in the matrix Indicates o i and p j The cosine similarity of . It is expressed as follows:
[0083]
[0084] For S t In the scene text dimension, the maximum similarity between each word and the image description text is obtained to obtain a vector of length k. Then, the vector is averaged to obtain the word-level similarity S between the scene text and the image description text. w , expressed as:
[0085]
[0086] Step S5-3, the input picture and image description text, according to step S3 and step S4, the word features O and P of the scene text and the image description text are obtained, and the attention similarity between the two texts is calculated by using the stacked cross attention mechanism. First, the cosine similarity between the scene text words and the image description text words is calculated, which is expressed as:
[0087]
[0088] where s i,j It is expressed as the similarity between the i-th scene text word and the j-th image description text word. Then the similarity score is normalized and expressed as:
[0089]
[0090] In order to use the attention mechanism between the scene text and the image description text, the image description text words are weighted to obtain a new image description text feature representation, which is expressed as:
[0091]
[0092] where α i,j It is expressed as:
[0093]
[0094] Where λ is the temperature coefficient, α i,j Represents the attention weight during the attention point multiplication operation. After attention weighting The correlation between the words in the scene text is expressed by the cosine similarity between the two:
[0095]
[0096] Finally, all word similarities are averaged to obtain the similarity S calculated by the stacked attention mechanism. a , expressed as:
[0097]
[0098] S5-4, the three similarities are weighted to obtain the final similarity S between the scene text and the image description sum , and the corresponding second sub-model is obtained:
[0099] S sum =S s +λ 1 S w +λ 2 S a
[0100] where λ 1 ,λ 2 is the balancing parameter, where λ 1 =4,λ 2 =6.
[0101] Step S6, integrating the first sub-model of step S2 and the second sub-model of step S5 into a unified framework to obtain a scene text retrieval model based on text detection and semantic matching;
[0102] The step S6 further comprises the following steps:
[0103] Step S6-1, integrate the model of step S2 and the model of S5 into a unified framework to obtain a scene text cross-modal retrieval model based on text detection and semantic matching.
[0104] S=S c +λ 3 S sum
[0105] Among them, λ 3 is the balancing parameter, λ 3 The value is 1. c is the cross-modal retrieval similarity obtained in step S2, S sum It is the sentence level, word level and similarity obtained by step S5 using the stacked cross attention mechanism.
[0106] The method of using the above model for cross-modal retrieval is as follows:
[0107] First, load the trained neural network parameters into the model. For image description text retrieval using images, the image to be retrieved and the image description text in the dataset are passed into the model as input. The model calculates the similarity between the image and other image description texts and outputs the top ten image description texts that are most similar to the input image.
[0108] For image retrieval using image description text, the image description text to be retrieved and the images in the dataset are passed into the model. The model calculates the similarity between the image description text and other images and outputs the top ten images most similar to the input image description text.
[0109] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. All equivalent methods or changes that do not deviate from the technical creation of the present invention should be included in the scope of protection of the present invention.
Claims
1. A scene text retrieval model based on text detection and semantic matching, It is characterized in that The model is obtained through the following steps: S1, extract the regional features of the image and the word-level features of the image description text, and map the two features to a common semantic space through a multi-layer perceptron to obtain image features V and text features E; S2, using cosine similarity to calculate the similarity between the two, the model is optimized and trained through the triple loss function, and finally the cross-modal retrieval similarity S is obtained c ; S3, extract the text information in the image, namely the scene text, and use fastText to extract the word features O of the scene text; S4, use StanfordCoreNlp to process the image description text, select the words that meet the semantic requirements and extract the word features P of the image description text through fastText; S5, for different levels of text features, uses word and sentence level features to calculate similarity, and uses the stacked cross attention mechanism to calculate, so that the model models the semantic relationship between scene text and image description text, and the three similarities are weighted to obtain the final similarity S between scene text and image description text. sum ; The specific implementation of S5 includes: S5-1, the input image and image description text are obtained according to step S3 and step S4. The word features O and P of the scene text and the image description text are averaged and fused to obtain the sentence level representation O of the current scene text. s : The word features of the image description text are averaged and fused to obtain the sentence level representation P of the current image description text. s : The cosine similarity is used to calculate the sentence-level similarity between the scene text and the image description text, and S s Represents the sentence-level similarity, expressed as: S5-2, the input image and image description text are used to obtain the word features O and P of the scene text and the image description text according to step S3 and step S4, and the word features O and P of the scene text and the image description text are formed by splicing the word features, which are represented as O and P respectively. w =[o i ,...,o k ], P w =[p i ,...,p l ],O w and P w The sizes are k*300 and l*300 respectively; construct the k*l cosine similarity matrix S t , each value in the matrix Indicates o i and p j The cosine similarity is expressed as follows: For S t In the scene text dimension, the maximum similarity between each word and the image description text is obtained to obtain a vector of length k. Then, the vector is averaged to obtain the word-level similarity S between the scene text and the image description text. w , expressed as: S5-3, the input picture and image description text, according to step S3 and step S4, obtain the word features O and P of the scene text and the image description text, and calculate the attention similarity between the two by using the stacked cross attention mechanism; the details are as follows: First, the cosine similarity between the scene text words and the image description text words is calculated, expressed as: where s i,j Represented as the similarity between the i-th scene text word and the j-th image description text word; Then the obtained cosine similarity is normalized and expressed as: Perform weighted operations on the words in the image description text to obtain a new image description text feature representation, which is expressed as: where α i,j It is expressed as: Where λ is the temperature coefficient, α i,j Represents the attention weight during the attention point multiplication operation; After attention weighting The correlation between the words in the scene text is expressed by the cosine similarity between the two: Finally, all word similarities are averaged to obtain the similarity S calculated by the stacked attention mechanism. a , expressed as: S5-4, the three similarities are weighted to obtain the final similarity S between the scene text and the image description text sum : S sum =S s +λ 1 S w +λ 2 S a where λ 1 ,λ 2 is the balancing parameter; S6, integrates S2 and S5 to obtain a scene text retrieval model based on text detection and semantic matching.
2. A scene text retrieval model based on text detection and semantic matching according to claim 1, It is characterized in that The specific implementation of S1 includes: S1-1, given an image I, use the pre-trained FasterRCNN to detect n regions of interest r in the image i , and extract the corresponding regional features f i ; Then use the multi-layer perceptron to transform the image region features f i Mapped to the common feature space to get v i : v i =MLP v (f i ) Among them, MLP v Represents the multi-layer perceptron corresponding to the image, and the obtained image features are expressed as S1-2, given a sentence T, for the i-th word in the sentence, use the one-hot encoding w i Indicates the position of the word in the vocabulary, using the mapping matrix W e w i Mapped into a 300-dimensional vector, represented as x i =W e w i ,i∈[1,m], where m represents the number of words in the sentence, and Bi-GRU is used to convert x i Mapped to word features; Bi-GRU includes a forward GRU, from w 1 To w m Read sentence T as follows: and a backward GRU, from w m To w 1 Read sentence T as follows: The final word feature e i By and Take the average method to fuse, so that the word feature fusion w i The context information of the surrounding sentences is expressed as i∈[1,m]; then use a multi-layer perceptron to map the image description to a common feature space to obtain e i : e i =MLP e (f i ) Among them, MLP e Represents the multi-layer perceptron corresponding to the image, and the obtained text feature is expressed as 3. A scene text retrieval model based on text detection and semantic matching according to claim 1, It is characterized in that The specific implementation of S2 includes: Use cosine similarity to calculate the similarity between image feature V and text feature E to obtain the similarity result of cross-modal retrieval; Use image aggregators and text aggregators visual (·) and f text (·) to aggregate and embed the image feature V and text feature E to obtain the aggregated features α, β: The similarity between image I and image description T is calculated by cosine similarity, and S c (v, e) represents the similarity of cross-modal retrieval, expressed as: Train the first sub-model of S2 using triplet loss: Where Δ is a hyperparameter and (v,e) represents the dataset The positive sample pairs in represents the hardest negative sample of v, represents the hardest negative sample of t, [x] + ≡max(0,x), using the ternary ranking loss to shorten the distance between positive sample pairs, where t ′ and v ′ is an intermediate variable.
4. A scene text retrieval model based on text detection and semantic matching according to claim 1, It is characterized in that The specific implementation of S3 includes: S3-1, for a given image I, input the image into the Rosetta OCR image text extraction system to extract all OCR tokens in the image, where OCR tokens are words recognized from the image; for each input image, use OCR to extract word text, i.e., scene text; S3-2, preprocessing the extracted scene text, first cleaning the data, deleting symbols, single characters and other recognized text; S3-3, perform part-of-speech analysis and screening on the scene text obtained in step S3-2, and send the scene text to StanfordCoreNlp for semantic analysis. The image descriptions and Rosetta OCR scene texts in the dataset are analyzed. The image descriptions and scene texts have a large number of words with the same part of speech, which correspond to the part-of-speech tags defined in StanfordCoreNlp, namely NN (noun, common, singular or numerous), NNS (noun, common, plural), NNP (noun, proper, singular), CD (number, cardinality), and JJ (adjective or numeral, ordinal). The scene texts preprocessed by S3-2 are screened for part of speech, and words with the part of speech included in the above five parts of speech are selected. Finally, k scene texts are selected for subsequent tasks. The final scene text is extracted into a 300-dimensional feature vector through the pre-trained fastText model. The word features of the scene text obtained by fastText are represented as follows:
5. A scene text retrieval model based on text detection and semantic matching according to claim 1, It is characterized in that The specific implementation of S4 includes: S4-1, perform part-of-speech analysis on the image description text, i.e., the unprocessed word text, and send the image description text to StanfordCoreNlp for semantic analysis to obtain the part-of-speech tag corresponding to each word; S4-2, by The image descriptions and Rosetta OCR scene texts in the image description text are analyzed, and the image description words with the parts of speech of NN, NNS, NNP, CD, and JJ are selected from the image description text; S4-3, the final image description text is extracted into a 300-dimensional feature vector through the pre-trained fastText model. The word feature representation of the input image description text obtained by fastText is:
6. A scene text retrieval model based on text detection and semantic matching according to claim 1, It is characterized in that The fusion method of the scene text retrieval model based on text detection and semantic matching of S6 is: S=S c +λ 3 S sum Among them, λ 3 is the balancing parameter, S c is the cross-modal retrieval similarity obtained in step S2, S sum It is the sentence level, word level and stacked cross attention similarity obtained by step S5.
7. A cross-modal retrieval method of a scene text retrieval model based on text detection and semantic matching according to any one of claims 1 to 6, It is characterized in that The cross-modal retrieval method is specifically as follows: First, the trained neural network parameters are loaded into the model. For image description text retrieval using images, the image to be retrieved and the image description text in the dataset are passed into the model as input. The model calculates the similarity between the image and other image description texts and outputs the top ten image description texts that are most similar to the input image. For image retrieval using image description text, the image description text to be retrieved and the images in the dataset are passed into the model. The model calculates the similarity between the image description text and other images and outputs the top ten images most similar to the input image description text.
8. A computer device, It is characterized in that The computer device has built-in execution instruction code or stored program code of the scene text retrieval model based on text detection and semantic matching as described in any one of claims 1 to 6, or execution instruction code or stored program code of the cross-modal retrieval method as described in claim 7.
Citation Information
Patent Citations
SLAM loopback detection method combined with scene text semantic information
CN111767854A
Cross-modal image text retrieval method of hybrid fusion model
CN112784092A