Deep Learning Voice-Assisted Character Recognition Method and Device Based on Pluggable Modules
By introducing a deep learning voice assistance method with pluggable modules into the scene text recognition network, using voice information for cross-modal interaction, the problem of editing errors in traditional methods is solved, the recognition accuracy is improved and the inference speed is maintained.
Patent Information
- Application Number
- CN202310111405.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-02-07
AI Technical Summary
Existing scene text recognition methods are not effective when dealing with editing errors, and traditional visual processing and language modeling processes are difficult to effectively solve the addition, modification or replacement errors.
Deep learning voice-assisted text recognition method based on pluggable modules is adopted. By using speech synthesis tools to generate speech data, combining image features and speech features, the voice decoder in the pluggable module is used to perform cross-modal interaction, optimize the recognition network weight parameters, and improve recognition accuracy.
It significantly improves the accuracy of text recognition, solves the problem of editing errors, and does not affect the inference speed of the original recognition network, and does not require additional labeling of voice data.
Smart Images

Figure CN116434732B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and computer vision, and more specifically, relates to a deep learning voice-assisted text recognition method based on a pluggable module. Background Art
[0002] Text is a basic tool for human communication and plays a crucial role in human information understanding. In daily life, natural scene text is everywhere. Reading text from natural images is a long-standing and deeply thought-provoking problem in computer vision. Driven by deep neural networks, text recognition technology has made remarkable progress in the past decade, promoting the development of multiple applications such as document scanning, data archiving, and ancient book recognition.
[0003] Traditional scene text recognition methods usually adopt a process of visual processing and language modeling to process text data. The former uses deep feature extraction technology for sequential prediction, and the latter is used to correct the recognition results implicitly or explicitly. Although this kind of process has greatly promoted the development of the field, editing errors such as addition, modification, or replacement have still not been well solved. Summary of the Invention
[0004] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention proposes a deep learning voice-assisted text recognition method based on a pluggable module, which mines voice information that has not been taken seriously before and assists the traditional network in text recognition.
[0005] To achieve the above object, according to one aspect of the present invention, there is provided a deep learning voice-assisted text recognition method based on a pluggable module, the method comprising:
[0006] Step 1: Data acquisition, using a publicly available synthetic text dataset as image training data, extracting labels as a corpus, and using a speech synthesis tool to pair and generate a certain number of voice data;
[0007] Step 2: Feeding the image-voice data into a recognition network to obtain image features and voice features respectively;
[0008] Step 3: Feeding the image features into a recognition decoder to output a predicted character sequence;
[0009] Step 4: Connecting the pluggable module to the scene text recognition network, and obtaining spectral features from the image features and voice features through a voice decoder in the pluggable module;
[0010] Step 5: The recognition network calculates the recognition loss, the pluggable module calculates the voice spectrum loss, and backpropagation optimizes the weight parameters of the recognition network;
[0011] Step 6: In the inference stage, pull out the pluggable module, and let the recognition network complete the recognition of the scene text image, and return the recognition result.
[0012] In one embodiment of the present invention, for the certain number of paired text-speech data sets synthesized in the above Step 1, the text and speech need to correspond one by one. It is not necessary to cover all the speech data. The subsequent network will automatically determine whether the input picture carries speech information. There is no restriction on the speech synthesis tool, but it is necessary to ensure that the input is text and the output is audio information, so as to construct the picture-speech training pair.
[0013] In one embodiment of the present invention, in the above Step 1: Through Fourier transform, the speech information is transformed from the time domain to the frequency domain, and the formula is: where s(n) represents the input signal, the original signal is divided into N small windows, and the fast Fourier transform is performed on the signal for each small window. W(n) is the sliding window function, and the j in the exponential part represents the imaginary unit. S(f, k) is a function of frequency f and time step k; the short-time Fourier transform is used to obtain the frequency domain feature S(f, k) of the speech signal, and it is converted from the linear scale to the Mel scale.
[0014] In one embodiment of the present invention, the formula for the conversion method from the linear scale to the Mel scale is: where f represents the frequency magnitude on the linear scale, and m represents the frequency magnitude on the Mel scale. Substitute f in the formula with m for S(f, k), and the representation S(m, k) of the Mel spectrum is obtained. The finally obtained Mel spectrum rather than the original speech signal is used as the supervision information.
[0015] In one embodiment of the present invention, the above Step 2 specifically includes: Given a color RGB text picture X with a height of H, a width of W, and a color channel number of 3, Θ enc represents the parameters of the image encoder. After the picture passes through the image encoder, the feature vector of the picture is obtained The specific formula is: The feature vector will be used as the input of the subsequent recognition decoder and speech decoder.
[0016] In one embodiment of the present invention, the above Step 3 specifically includes: sending the feature vector generated by the image encoder in Step 2 into the recognition decoder, and the recognition decoder will generate a predicted character sequence where n is the length of the prediction sequence, and y i represents the i-th predicted character. The specific formula is where RecoDecoder represents the recognition network, and Θ rec are the parameters of the recognition network.
[0017] In one embodiment of the present invention, the speech decoder in step four is composed of a pre-network Prenet, a visual-speech decoder VADecoder, and a Mel spectrum linear layer MelLinear; the Mel spectrum Z is input into the pre-network Prenet. Z is a vector of T×80 dimensions, where T represents the length of the audio in the time dimension. The pre-network consists of two fully connected layers and a ReLU activation layer. Each fully connected layer uses 256 neurons. The role of the pre-network layer is to project the spectrum into the phoneme subspace, so that the subsequent decoding network can better capture the relationship between the scene text image and the speech information. The specific formula is: is the transformed audio feature Θ pre represents the parameters of the pre-network.
[0018] In one embodiment of the present invention, the visual-speech decoder VADecoder is implemented based on Transformer, and uses the Mel spectrum feature as the query, and the image feature as the key and value to achieve cross-modal interaction between the picture and the speech. The specific formula is where represents the decoded audio vector; the Mel spectrum linear layer is a linear fully connected layer, which is used to convert the output result of the decoder into the Mel spectrum feature. The formula is:
[0019] In one embodiment of the present invention, step five specifically includes: the cost function is composed of an identification loss function and a spectrum prediction loss function. The formula is: L = L rec +α·L mel , where L rec is the identification loss of the identification decoder in the general scene text recognition network, and L mel is the loss of the speech decoder in the pluggable module.
[0020] According to another aspect of the present invention, there is also provided a deep learning speech-assisted text recognition device based on a pluggable module, including at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to complete the deep learning speech-assisted text recognition method based on the pluggable module according to any one of claims 1-9.
[0021] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects are obtained:
[0022] The present invention uses a plug-and-play module structure, which is convenient for reuse in the current open-source scene text recognition methods and can be unplugged during the inference and prediction stage, without affecting the inference speed of the original recognition network while significantly improving the recognition performance. Moreover, the present invention does not require additional speech data annotation and only needs an existing speech synthesis engine to obtain speech data for free. The addition of speech data can solve the editing and erasing errors to a certain extent due to the rationality of pronunciation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a schematic flow chart of a deep learning speech-assisted text recognition method based on a pluggable module in an embodiment of the present invention;
[0024] Figure 2 FIG. is a schematic flow chart of text-to-speech conversion in an embodiment of the present invention;
[0025] Figure 3 FIG. is a schematic structural diagram of a speech decoder in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0027] In order to solve the problems existing in the prior art, the present invention proposes a deep learning speech-assisted text recognition method based on a pluggable module. Our goal is to enable the text detection and recognition network to not only have the ability to "see", but also have the ability to "hear", that is, the ability to identify and predict audio information. As Figure 1 shown, Figure 1 In the recognition network shown, the solid line part is a general scene text recognition method, and the dotted line part is the deep learning speech-assisted text recognition part based on the pluggable module innovatively proposed by the present invention. The recognition network mainly consists of three core parts: an image encoder, a recognition decoder, and a speech decoder (i.e., the main part of the pluggable module of the present invention. The pluggable module of the present invention also includes the part for calculating loss mentioned in step five below). The image encoder and the recognition decoder are the core parts of the general scene text recognition method, and the speech decoder is the core part of the deep learning speech-assisted text recognition method based on the pluggable module proposed by the present invention.
[0028] As Figure 1 shown, the deep learning speech-assisted text recognition method based on the pluggable module proposed by the present invention includes the following steps:
[0029] Step 1: Data acquisition. Use a publicly available synthetic text dataset as the image training data, extract the labels as the corpus, and use a speech synthesis tool to pair and generate a certain amount of speech data.
[0030] Specifically, for the certain amount of paired image-speech datasets synthesized in Step 1, the images and speech need to correspond one by one. Considering the synthesis efficiency of speech, it is not necessary for all speech data to cover all images, that is, not every image necessarily has its corresponding speech data, but every piece of speech data must have its corresponding image data. Experiments have shown that only 20% of the speech data is sufficient to significantly improve the recognition accuracy. Subsequently, the network will automatically determine whether the input image carries speech information and perform corresponding processing. For details, please refer to Step 5. The speech synthesis tool is not restricted, and open-source speech synthesis tools (such as PaddleOCR developed by Baidu PaddlePaddle) can be used. It is necessary to ensure that the input is text and the output is audio information.
[0031] Construct image-speech training pairs;
[0032] Specifically, the sampling rate of the speech synthesis tool is set to 22050 here, which means sampling 22050 times per second. The higher the sampling rate, the higher the quality of the obtained speech information, but at the same time, it will also increase the generation time and occupy more memory. The synthesized speech data is saved in the format of a wav file. By visualizing the wav file, the waveform of the speech can be obtained, as Figure 2 shown. Through the Fourier transform, the speech information can be transformed from the time domain to the frequency domain. Considering that the frequency domain characteristics of the signal change with time, the short-time Fourier transform is used here, as shown in formula (1):
[0033]
[0034] where s(n) represents the input signal. The original signal is divided into N small windows, and the fast Fourier transform is performed on the signal in each small window. W(n) is the sliding window function. Here, the step size of the sliding window is set to 275.625 milliseconds, and the length of the sliding window is 1102.5 ms. The step size and length of the sliding window will affect the number of times the Fourier transform is calculated on the original signal. The exponent part j represents the imaginary unit, and S(f, k) is a function of frequency f and time step k.
[0035] Use the short-time Fourier transform to obtain the frequency domain characteristics S(f, k) of the speech signal. Since humans' perception of frequency is not linear, and humans are more sensitive to low-frequency signals than high-frequency signals. To be closer to human perception and make the learning of the recognition network more effective, the Mel scale commonly used in the scientific community is used here. There are multiple conversion methods from the linear scale to the Mel scale, and a commonly used conversion method is shown in formula (2):
[0036]
[0037] Among them, f represents the frequency magnitude on a linear scale, and m represents the frequency magnitude on a Mel scale.
[0038] By replacing f in S(f, k) in formula (1) with m, the representation of the Mel spectrum S(m, k) is obtained, and the visualization result is as Figure 2 shown. Our method uses the finally obtained Mel spectrum rather than the original speech signal as the supervision information.
[0039] Step 2: Feed the image-speech data into the recognition network to obtain image features and speech features respectively;
[0040] Specifically, the speech data is converted into a Mel spectrum in the manner described in Step 1, and the Mel spectrum will serve as the supervision signal, i.e., the speech feature, of the pluggable module.
[0041] Specifically, the image data is fed into the image encoder to obtain image features. Given a color RGB text image X with a height of H, a width of W, and 3 color channels (usually a batch containing multiple images during specific training, and here only the case of one image is considered for simplicity), Θ enc represents the parameters of the image encoder, and the feature vector of the image is obtained after the image passes through the image encoder Feature vector will be used as the input for the subsequent recognition decoder and speech decoder. The specific formula is as shown in (3):
[0042]
[0043] The image encoder can be flexibly initialized to different types according to different recognition networks, such as CNN (Convolutional Neural Network), ResNet (Residual Neural Network), and various models based on Transformer, which is convenient for modification and adjustment;
[0044] Here, the images input to the image encoder come from synthetic datasets, usually the commonly used SynText, MJSynth, and SynAdd datasets in academia and industry, or real-scene datasets and multilingual datasets can also be used.
[0045] Step 3: Feed the image features into the recognition decoder to output the predicted character sequence;
[0046] Specifically, the feature vector generated by the image encoder in Step 2 is fed into the recognition decoder, and the recognition decoder will generate the predicted character sequence where n is the length of the predicted sequence, and y iRepresents the predicted i-th character. The specific formula is shown in (4). RecoDecoder represents the recognition network, and Θ rec are the parameters of the recognition network:
[0047]
[0048] In real application scenarios, there are three commonly used recognition networks that can be used as recognition decoders, Connectionist Temporal Classification (CTC), attention-based sequence prediction model (Attn), and Transformer-based sequence prediction model (Trans).
[0049] Step four: Connect the pluggable module to the scene text recognition network. The image features and speech features pass through the speech decoder in the pluggable module to obtain spectral features;
[0050] The speech decoder described in step four is the core part of the pluggable module. The structure of the speech decoder is as Figure 3 shown, consisting of a pre-network Prenet, a visual-audio decoder VADecoder, and a Mel spectrogram linear layer MelLinear;
[0051] Specifically, input the Mel spectrogram Z into the pre-network Prenet. Z is a vector of T×80 dimensions, where T represents the length of the audio in the time dimension. The pre-network consists of two fully connected layers and a ReLU activation layer, and each fully connected layer uses 256 neurons. The role of the pre-network layer is to project the spectrogram into the phoneme subspace, enabling the subsequent decoding network to better capture the relationship between the scene text image and speech information. The specific formula is expressed as (5):
[0052]
[0053] is the transformed audio feature Θ pre represents the parameters of the pre-network.
[0054] Specifically, the visual-audio decoder VADecoder is implemented based on Transformer, using the Mel spectrogram feature as the query, and the image feature as the key and value to achieve cross-modal interaction between the picture and speech. The specific formula is expressed as (6):
[0055]
[0056] where Denote the decoded audio vector. In specific implementation, the audio decoder includes 3 identical decoder layers, and each decoder layer includes a masked multi-head attention layer, a multi-head attention layer, and a feed-forward network.
[0057] Specifically, the Mel spectrogram linear layer is a linear fully-connected layer, which is used to convert the output result of the decoder into Mel spectrogram features for calculating the loss, as shown in formula (7) specifically:
[0058]
[0059] Step Five: The recognition network calculates the recognition loss, the pluggable module calculates the speech spectrogram loss, and backpropagation optimizes the weight parameters of the recognition network;
[0060] Specifically, the cost function in Step Three is composed of a recognition loss function and a spectrogram prediction loss function, as shown in formula (8) specifically:
[0061] L = L rec + α·L mel (8)
[0062] Among them, L rec is the recognition loss of the recognition decoder in the general scenario text recognition network, and L mel is the loss of the speech decoder in the pluggable module. Here, the L1 loss is used in both cases. α is the weight coefficient of the speech decoder loss, and the default setting is 1.0.
[0063] Specifically, as described in Step One, the speech data does not necessarily cover the picture data completely. Therefore, the recognition network needs to determine whether the input training data contains speech data. If it does not contain speech data, then L mel , L mel = 0.
[0064] Here, the backpropagation algorithm is specifically used to calculate the gradient of the obtained loss and let it backpropagate in the network to optimize the recognition network parameters.
[0065] Step Six: In the inference stage, the pluggable module is pulled out, and the recognition network completes the recognition of the scene text image and returns the recognition result.
[0066] In the above steps, in the inference and prediction stage, the pluggable module is pulled out. Therefore, the pluggable module will not affect the inference speed, enabling the original recognition network to improve the recognition accuracy while maintaining the speed.
[0067] In the inference and prediction stage, the pluggable module is unplugged, and only images are used during inference, which also conforms to the usual recognition logic. Currently, the vast majority of scene text datasets do not contain voice data. The pluggable module will not affect the inference speed, enabling the original recognition network to improve the recognition accuracy while maintaining the speed.
[0068] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A deep learning voice-assisted text recognition method based on a pluggable module, characterized in that The method includes: Step 1: Data acquisition. Use a publicly available synthetic text dataset as image training data, extract the labels as a corpus, and use a speech synthesis tool to pair and generate a certain number of speech data. Step 2: Send the image-speech data into the recognition network to obtain image features and speech features respectively. Step 3: Send the image features into the recognition decoder to output the predicted character sequence. Step 4: Connect the pluggable module to the scene text recognition network. The image features and speech features are passed through the speech decoder in the pluggable module to obtain spectral features. The speech decoder in Step 4 consists of a front-end network , a vision-speech decoder , and a mel-spectrogram linear layer . Input the mel-spectrogram into the front-end network . is a -dimensional vector, and T represents the length of the audio in the time dimension. The front-end network consists of two fully connected layers and a ReLU activation layer. Each fully connected layer uses 256 neurons. The role of the front-end network layer is to project the spectrum into the phoneme subspace, enabling the subsequent decoding network to better capture the relationship between the scene text image and speech information. The specific formula is: , is the transformed audio feature represents the parameters of the front-end network. The vision-speech decoder is implemented based on Transformer. It uses the mel-spectrogram feature as the query, and the image feature as the key and value to achieve cross-modal interaction between the picture and speech. The specific formula is , where represents the decoded audio vector. The mel-spectrogram linear layer is a linear fully connected layer used to convert the output result of the decoder into mel-spectrogram features. The formula is: ; Step 5: The recognition network calculates the recognition loss, and the pluggable module calculates the speech spectrum loss, and backpropagation is used to optimize the weight parameters of the recognition network. Step 6: In the inference stage, pull out the pluggable module, and the recognition network completes the recognition of the scene text image and returns the recognition result.
2. The deep learning voice-assisted text recognition method based on a pluggable module according to claim 1, wherein: For the certain number of paired text-speech datasets synthesized in Step 1, the text and speech need to correspond one by one. It is not necessary for all speech data to be covered. The subsequent network will automatically determine whether the input image carries speech information. The speech synthesis tool is not restricted, but it is necessary to ensure that the input is text and the output is audio information to construct image-speech training pairs.
3. The deep learning voice-assisted character recognition method based on a pluggable module according to claim 2, wherein In Step 1: Through Fourier transform, convert the speech information from the time domain to the frequency domain. The formula is: , where represents the input signal, the original signal is divided into N small windows, and the fast Fourier transform is performed on the signal for each small window. is the sliding window function, and the in the exponential part represents the imaginary unit. is the frequency and is a function of the time step k; the frequency domain characteristics of the speech signal are obtained using the short-time Fourier transform , and are converted from a linear scale to a mel scale.
4. A deep learning voice-assisted text recognition method based on a pluggable module according to claim 3, characterized in that, The formula for the conversion method from the linear scale to the Mel scale is as follows: where represents the frequency magnitude on the linear scale, represents the frequency magnitude on the Mel scale. Substituting in the formula with using we obtain the representation of the Mel spectrum and use the finally obtained Mel spectrum rather than the original speech signal as the supervision information.
5. A deep learning voice-assisted character recognition method based on a pluggable module according to claim 1 or 2, characterized in that Step 2 specifically includes: Given a color RGB text image with a height of H, a width of W, and 3 color channels , denotes the parameters of the image encoder. After passing through the image encoder, the feature vector of the image is obtained , and the specific formula is as follows: Feature vector Will be used as the input for subsequent recognition decoders and speech decoders.
6. A deep learning voice-assisted text recognition method based on a pluggable module according to claim 1 or 2, characterized in that Step 3 specifically includes: The feature vectors generated by the image encoder in step two are fed into the recognition decoder, and the recognition decoder will generate a predicted character sequence , where n is the length of the prediction sequence, represents the i-th predicted character, and the specific formula is , where r represents the recognition network, are the parameters of the recognition network.
7. A deep learning voice-assisted text recognition method based on a pluggable module according to claim 1 or 2, characterized in that Step 5 specifically includes: The cost function consists of an identification loss function and a spectrum prediction loss function, and the formula is: , where is the identification loss of the identification decoder in the general scene text recognition network, is the loss of the speech decoder in the pluggable module.
8. A deep learning voice-assisted text recognition device based on a pluggable module, characterized in that: It includes at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions that can be executed by the at least one processor. After the instructions are executed by the processor, they are used to complete the deep learning speech-assisted text recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Scene text recognition method based on semantic relevancy prediction and attention decoding
CN110717336A
Multi-modal semantic recognition service access method based on artificial intelligence
CN112201228A