A Handwritten Text Recognition Method and System Based on Sound Signals
Through the handwritten text recognition method and system based on sound signals, the random forest model and KNN classification algorithm are used to solve the problems of poor portability, high power consumption and privacy of smart wearable devices, and high accuracy handwritten text input in different environments is achieved, which is suitable for smart watches and other smart devices.
Patent Information
- Application Number
- CN202011504052.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-12-17
AI Technical Summary
The handwritten text input methods of existing smart wearable devices have problems such as poor portability, high power consumption, privacy issues, and low recognition accuracy in noisy environments, especially the obvious shortcomings in voice recognition and camera recognition.
Through the handwritten text recognition method based on sound signals, the random forest model and KNN classification algorithm are used, combined with sound signal acquisition and processing technology, contactless handwritten text input is realized, and a recognition system including a control module, a local sound signal acquisition module and a cloud computing module is built to adapt to the recognition of different environments and languages.
It realizes low-cost, low-power, and easy-to-port handwritten text input under the premise of ensuring user privacy. It has a wide range of application and high recognition accuracy. It can adapt to different environments and reduce the possibility of misidentification, and has a good user experience.
Smart Images

Figure CN114648772B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of handwritten text recognition, and in particular to a handwritten text recognition method and system based on sound signals. Background Art
[0002] In recent years, with the gradual popularization of consumption upgrading and technologies such as AI, VR, and AR, the functions of intelligent devices have become more and more diverse. Wearable intelligent devices such as smart watches and smart wristbands are loved by many users because of their portable and practical features. However, smart watches or other intelligent wearable devices are limited by a smaller interaction interface and a smaller number of interaction interfaces, resulting in poor interactivity with users. When users want to type or perform text input, it is not convenient enough, and handwritten text input is easily misrecognized, resulting in a significant reduction in the usability of some intelligent wearable devices. To solve this dilemma, many device manufacturers design the watch screen to be larger to provide more space to facilitate users to input text through buttons or touchscreens. However, the size of the screen is limited, and this approach also reduces the portability of the device. Some smart watches come with a stylus to facilitate user operations, but this additional hardware also makes the device inconvenient to use.
[0003] In order to enable users to perform text input on the device, many devices have added a voice recognition function to convert sound signals into text. Users can interact with the device directly without touching the device through voice interaction. However, voice recognition also has some deficiencies in some aspects. First, voice recognition is greatly affected by the environment, and the accuracy of voice recognition is relatively low in a noisy environment. Second, using voice recognition for text input is likely to involve user privacy issues, and some users who value personal privacy do not like to use voice interaction for text input. There are also many devices that collect the hand movements of users through a camera and use relevant image recognition algorithms to recognize the movements of users writing in the air to identify characters, so as to obtain the text input by the users. However, the power consumption of the camera is large, and the camera is easily restricted by the usage environment, and the effect is not good at night or in a relatively dim environment. At the same time, the camera also touches on user privacy.
[0004] Chinese Patent CN201911122436.7 discloses a handwritten text recognition method that outputs recognized characters by collecting the motion signals of the user's handwritten actions. However, this method requires a stylus accessory, and uses the motion information such as displacement and rotation collected by the sensors on the stylus accessory as the motion signal for handwritten text recognition. In the field of intelligent device interaction, it obviously reduces the portability of the device. Therefore, there is an urgent need for an intelligent device interaction system with low cost, low power consumption, easy to carry, and not involving user privacy to realize user text input. Summary of the Invention
[0005] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and provide a method and system for handwritten text recognition based on voice signals. On the premise of ensuring user privacy, the handwritten text input by the user is recognized through voice signals, realizing contactless handwritten text input recognition, which can recognize different languages in different environments, has a wide range of applications and high reliability.
[0006] The object of the present invention can be achieved by the following technical solutions:
[0007] A method for handwritten text recognition based on voice signals, comprising the following steps:
[0008] S1: Set the text recognition range W, where the text recognition range W includes M (M>0) different characters, W = {w1, w2,..., w M}; respectively generate preliminary recognition sets for the characters w1, w2,..., w M : I w1 , I w2 ,... I wM , where I wi (i>1) stores: a label indicating whether the character w i is stably recognized, a characteristic voice signal, and a classification data set corresponding one-to-one with the characteristic voice signal. The preliminary recognition result of the characteristic voice signal is w i ; initialize all M preliminary recognition sets to empty sets, where the labels are all initialized as not stably recognized;
[0009] S2: Obtain a text picture set, where the text picture set includes M handwritten character picture sets: H w1 , H w2 ,... H wM , where H wi (i>1) is a handwritten character picture of the character w i . Each handwritten character picture set includes N (N>2) handwritten character pictures of different users; convert each handwritten character picture in the M handwritten character picture sets into a voice signal respectively to obtain M voice signal sets: Y w1 , Y w2 ,... Y wM , corresponding to the characters w1, w2,..., w M respectively; based on the M voice signal sets, construct a random forest model, where the input of the random forest model is the voice signal and the output is multiple model recognition results; enter the initialization stage;
[0010] S3: Collect the original sound signal when the user inputs handwritten text, and preprocess it to obtain a characteristic sound signal. The characteristic sound signal is a time-distance sequence, specifically the distance between the user's hand and the sound signal receiving device at different times. If all the characters within the text recognition range W can be stably recognized, the initialization phase ends and the usage phase begins. Execute step S7. Otherwise, continue the initialization phase and execute step S4;
[0011] S4: Input the characteristic sound signal into the random forest model to obtain multiple model recognition results. The user confirms the preliminary recognition result m (m ∈ W) corresponding to the characteristic sound signal from multiple model recognition results; Obtain the preliminary recognition set I of the character m m , according to the preliminary recognition set I m Judge whether the character m is stably recognized according to the labels in it. If so, output the preliminary recognition result m and execute step S3; If not, execute step S5;
[0012] S5: Write the characteristic sound signal into the preliminary recognition set I m ; According to all the characteristic sound signals in the preliminary recognition set I m Calculate the characteristic sound vector S user , obtain the sound signal set r corresponding to the character m m , calculate the comprehensive similarity between the characteristic sound vector S user and all the sound signals in the sound signal set Y m respectively, select the L (L > 1) sound signals with the largest comprehensive similarity in the sound signal set Y m as the latest classification data set of the character m, and write the latest classification data set into the preliminary recognition set I m and correspond to the characteristic sound signal;
[0013] S6: Judge whether the preliminary recognition set I m reaches a stable state. If it reaches a stable state, update the label of the preliminary recognition set I m to be stably recognizable, that is, the character m can be stably recognized. Take the latest classification data set as the enhanced recognition set of the character m, output the preliminary recognition result m, and execute step S3; If it does not reach a stable state, directly output the preliminary recognition result m and execute step S3;
[0014] S7: Obtain the enhanced recognition sets of the characters w1, w2,..., w M , based on the KNN classification algorithm, obtain the sound signals in the M enhanced recognition sets that are closest to the characteristic sound signal as similar sound signals, and take the corresponding characters of the similar sound signals as the recognition results;
[0015] S8: Repeat step S3 until the handwritten text recognition ends, perform error correction on all the obtained recognition results, and obtain the final recognized text.
[0016] Further, step S2 includes the following steps:
[0017] S21: Obtain multiple handwritten text images of a user, perform stretching and scaling operations on them to unify the specifications of the handwritten text images, and use the text in each handwritten text image as the corresponding text of the handwritten text image;
[0018] S22: Traverse all the handwritten text images of this user, and put the handwritten text images with the corresponding text w1 into the handwritten text image set H of the text w1 w1 and put the handwritten text images with the corresponding text w2 into the handwritten text image set H of the text w2 w2 and so on, and put the handwritten text images with the corresponding text w M into the handwritten text image set H of the text w M ; wM ;
[0019] S23: Repeat step S21 until each handwritten text image set stores handwritten text images of N different users respectively, and obtain a text image set including M handwritten text image sets;
[0020] S24: Select a handwritten text image set from the text image set, and convert each handwritten text image in this handwritten text image set into a sound signal through the particle swarm algorithm to obtain a sound signal set converted from this handwritten text image set;
[0021] S25: Repeat step S24 until all handwritten text image sets are converted into sound signal sets, and obtain M sound signal sets: Y w1 , Y w2 and so on, Y wM , corresponding to the texts w1, w2,..., w M ;
[0022] S26: Build a random forest model based on the M sound signal sets. The input of the random forest model is the sound signal, and the output is multiple model recognition results;
[0023] S27: Enter the initialization stage.
[0024] Even further, step S24 includes the following steps:
[0025] S241: Select a handwritten text image set from the text image set;
[0026] S242: Obtain a handwritten text picture from the handwritten text picture set, perform a skeletonization operation on the handwritten text picture and traverse it to obtain the range, starting point, ending point, and number of strokes of the text strokes in the handwritten text picture;
[0027] S243: Generate a particle swarm. The initial position of each particle in the particle swarm is the starting point of the text stroke. According to the range and number of strokes of the text stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the velocity attributes of each particle, and the particles move, and the number of times is incremented by 1;
[0028] S244: Update the positions of the particles. If the position of a particle is not within the range of the text stroke, remove the particle from the particle swarm. According to the range and number of strokes of the text stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the velocity attributes of each particle, and the particles move, and the number of times is incremented by 1;
[0029] S245: If the number of times is equal to the number of strokes, execute step S246; otherwise, repeat step S244;
[0030] S246: Obtain the movement trajectories of all the particles in the current particle swarm and perform an averaging operation to obtain an average trajectory, and calculate the time-distance sequence of the average trajectory as the sound signal corresponding to the handwritten text picture;
[0031] S247: Repeat step S242 until all the handwritten text pictures in the handwritten text picture set are converted into sound signals, and obtain the sound signal set converted from the handwritten text picture set.
[0032] Further, the step S3 includes the following steps:
[0033] S31: Send a high-frequency sound signal to the human hand through a sound signal generating device, and receive the sound signal reflected by the human hand through a sound signal receiving device and use it as the original sound signal;
[0034] S32: Perform normalization processing, Gaussian filtering, and noise reduction processing on the original sound signal;
[0035] S33: Obtain the time-distance sequence as the characteristic sound signal;
[0036] S34: If all the texts within the text recognition range W can be stably recognized, the initialization stage ends, enter the usage stage, and execute step S7; otherwise, continue the initialization stage and execute step S4.
[0037] Further, the step S4 includes the following steps:
[0038] S41: Input the characteristic voice signal into the random forest model to obtain multiple model recognition results;
[0039] S42: The user finds the text entered in step S3 from the multiple model recognition results and confirms it as the preliminary recognition result m (m ∈ W) corresponding to the characteristic voice signal;
[0040] S43: Obtain the preliminary recognition set I of the text m m , and judge whether the text m is stably recognized according to the labels in the preliminary recognition set I m . If so, output the preliminary recognition result m and execute step S3; if not, execute step S5.
[0041] Further, step S5 includes the following steps:
[0042] S51: Write the characteristic voice signal into the preliminary recognition set I m ;
[0043] S52: Obtain all the characteristic voice signals in the preliminary recognition set I m and calculate the characteristic voice vector S user :
[0044]
[0045] where S j represents the characteristic voice signal in the preliminary recognition set I m , n represents the number of characteristic voice signals in the preliminary recognition set I m ;
[0046] S53: Obtain the voice signal set Y corresponding to the text m m , and calculate the Euclidean distance and Pearson similarity between the characteristic voice vector S user and all the voice signals in the voice signal set Y m respectively;
[0047] S54: Select a voice signal p (voice signal p ∈ Y m ) from the voice signal set Y m , and calculate the comprehensive similarity Distance between the characteristic voice vector S user and the voice signal p:
[0048]
[0049] where S user is the characteristic voice vector, Image p is the voice signal p, d user_p is the Euclidean distance between the characteristic voice vector and the voice signal p, max dis the maximum value of the Euclidean distances between the characteristic sound vector and all sound signals, min d is the minimum value of the Euclidean distances between the characteristic sound vector and all sound signals, Pearson(S user , Image p ) is the Pearson similarity between the characteristic sound vector and the sound signal p, and α and β are weight coefficients;
[0050] S55: Repeat step S54 until the characteristic sound vector S user and the comprehensive similarity between all sound signals in the sound signal set Y m are obtained;
[0051] S56: Select the L sound signals with the largest comprehensive similarity in the sound signal set Y m as the latest classification data set of the text m, and write the latest classification data set into the preliminary recognition set I m and correspond it to the characteristic sound signal.
[0052] Further, the step S6 includes the following steps:
[0053] S61: Obtain all classification data sets in the preliminary recognition set I m . If the number of classification data sets is equal to the preset recognition threshold F1 (F1 > 2), then the preliminary recognition set I m reaches a stable state, and step S64 is executed; if the number of classification data sets is greater than or equal to the preset judgment threshold Kp (F1 > Kp ≥ 2) and less than the preset recognition threshold F1, then step S62 is executed; if the number of classification data sets is less than the preset judgment threshold Kp, then step S65 is executed;
[0054] S62: Obtain the latest Kp classification data sets written in the preliminary recognition set I m ;
[0055] S63: Calculate the change error between the Kp classification data sets. If the change error is less than the preset stability threshold, then the preliminary recognition set I m reaches a stable state, and step S64 is executed. Otherwise, the preliminary recognition set I m does not reach a stable state, and step S65 is executed;
[0056] S64: Update the label of the preliminary recognition set I m to be stably recognized, that is, the text m can be stably recognized. Take the latest classification data set as the enhanced recognition set of the text m, output the preliminary recognition result m, and execute step S3;
[0057] S65: Output the preliminary recognition result m and execute step S3.
[0058] Further, in step S8, all the obtained recognition results are corrected at the word level using a text error correction data packet based on natural language processing to obtain the final recognized text.
[0059] Further, it further includes step S9, and the specific content of step S9 is: obtaining the feedback result of the user on the recognition accuracy of the handwritten text; if the feedback result is that the accuracy is lower than the preset accuracy threshold, then execute step S1, otherwise, execute step S3.
[0060] A handwritten text input recognition system based on voice signals, based on the above-mentioned handwritten text input recognition method, includes a control module, a local voice signal acquisition module, a local storage module, and a cloud computing module;
[0061] The control module, the local voice signal acquisition module, and the local storage module are integrated on the user side, and the cloud computing module is set in the cloud;
[0062] The control module is respectively connected to the local voice signal acquisition module and the local storage module, and the control module is communicatively connected to the cloud computing module;
[0063] The local voice signal acquisition module includes a voice signal generating device, a voice signal receiving device, and a data preprocessing unit. The voice signal generating device is used to send high-frequency voice signals to the human hand, and the voice signal receiving device is used to receive the voice signals reflected by the human hand and transmit them as the original voice signals to the data preprocessing unit; the data preprocessing unit is used to preprocess the original voice signals to obtain feature voice signals;
[0064] The local storage module is used to store the enhanced recognition set;
[0065] The cloud computing module includes a data calculation unit and a storage unit connected to each other; the data calculation unit is used for data calculation during the text recognition process; the storage unit is used to store the text picture set and the data during the text recognition process.
[0066] Compared with the prior art, the present invention has the following beneficial effects:
[0067] (1) On the premise of ensuring user privacy, the handwritten text input by the user is recognized through voice signals, realizing contactless handwritten text input recognition, which can recognize different languages in different environments, has a wide range of applications, and high reliability.
[0068] (2) Text recognition includes an initialization stage and a usage stage. In the initialization stage, based on a random forest model and user confirmation, the text recognition result is output. After the initialization stage ends, an enhanced recognition set of all the text within the text recognition range is obtained. After entering the usage stage, the KNN classification algorithm is directly used to find the voice signal in the enhanced recognition set that is closest to the characteristic voice signal of the user's handwritten text, thereby realizing text recognition, reducing the possibility of misrecognition, being able to recognize continuously input text, and can be used without prior training.
[0069] (3) Voice signal acquisition is performed based on the built-in voice signal generating device and voice signal receiving device on the intelligent device. First, the enhanced recognition set is obtained through cloud computing, and then text recognition is performed locally based on the enhanced recognition set and the KNN classification algorithm. No additional device is required, it is easy to carry, and the cost and power consumption are relatively low.
[0070] (4) A feedback module is provided. When the user believes that the recognition accuracy is low or the user's handwriting habit changes, learning can be restarted to obtain the enhanced recognition set again, and the user experience is high.
[0071] (5) Calculate the feature vector by combining the previously input characteristic voice signal, which is suitable for the user's handwriting habit, logical, calculate the comprehensive similarity by integrating the Euclidean distance and Pearson similarity, and can more objectively reflect the similarity, improving the recognition accuracy. Description of the Drawings
[0072] Figure 1 is the flowchart of handwritten text recognition;
[0073] Figure 2 is the architecture diagram of the handwritten text recognition system in the embodiment;
[0074] Figure 3 is the schematic diagram of the text recognition accuracy of different languages in the embodiment;
[0075] Figure 4 is the schematic diagram of the text recognition accuracy in different environments in the embodiment. Detailed Embodiment
[0076] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manner and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiments.
[0077] Embodiment 1:
[0078] A handwritten text input recognition system based on voice signals, based on the above handwritten text input recognition method, includes a control module, a local voice signal acquisition module, a local storage module, and a cloud computing module;
[0079] The control module, the local sound signal acquisition module, and the local storage module are integrated at the user side, and the cloud computing module is set up in the cloud;
[0080] The control module is respectively connected to the local sound signal acquisition module and the local storage module, and the control module is communicatively connected to the cloud computing module;
[0081] The local sound signal acquisition module includes a sound signal generating device, a sound signal receiving device, and a data preprocessing unit. The sound signal generating device is used to send high-frequency sound signals to the human hand, and the sound signal receiving device is used to receive the sound signals reflected by the human hand and transmit them as original sound signals to the data preprocessing unit; the data preprocessing unit is used to preprocess the original sound signals to obtain characteristic sound signals;
[0082] The local storage module is used to store the enhanced recognition set;
[0083] The cloud computing module includes a data calculation unit and a storage unit connected to each other; the data calculation unit is used for data calculation in the text recognition process; the storage unit is used to store the text picture set and the data in the text recognition process.
[0084] In this embodiment, the intelligent device takes a smart watch as an example, and the smart watch is equipped with a speaker and a microphone. When the user performs handwritten text input, the speaker acts as the sound signal generating device to send high-frequency signals, and the microphone acts as the sound signal receiving device to receive the sound signals reflected by the human hand. In the data preprocessing unit, the collected original sound signals are subjected to normalization processing, Gaussian filtering, and noise reduction processing to obtain characteristic sound signals, that is, time-distance sequences, indicating the distance between the user's hand and the microphone at any moment.
[0085] As Figure 2 shown, the user writes without contact on the intelligent device, and after collecting the original sound signals through the speaker and the microphone, they are transmitted to the data preprocessing unit for filtering and noise reduction to obtain characteristic sound signals.
[0086] When used for the first time, the local storage module does not have an enhanced recognition set, and it enters the initialization stage, and the characteristic sound signals are transmitted to the cloud computing module by means of wireless communication (such as WIFI communication, Bluetooth communication, etc.);
[0087] In the cloud, the storage unit stores a text data set, and the data calculation unit converts the text data set into a sound signal set and stores it in the storage unit, and constructs a random forest model based on the sound signal set;
[0088] After receiving the feature voice signal, the cloud computing module inputs it into the random forest model. The random forest model outputs multiple model recognition results. The user finds the text he / she input from the multiple model recognition results and confirms it to obtain the preliminary recognition result m, and outputs the preliminary recognition result.
[0089] In the initialization stage, every time the text input by the user passes through the random forest model to output multiple model recognition results first, and then is confirmed by the user. If the text m has not been stably recognized, the cloud computing calculates the comprehensive similarity to obtain the latest classification data set of the preliminary recognition result m, and updates the preliminary recognition set. If the preliminary recognition set reaches a stable state, it means that the text m can be stably recognized, and the enhanced recognition set of the text m can be obtained and written into the local storage module. If the text m has been stably recognized, there is already an enhanced recognition set of the text m in the local storage module, and the next text input recognition can be directly performed.
[0090] When the enhanced recognition sets of all texts in the text recognition range are found, the initialization stage ends.
[0091] After the initialization stage ends, the local storage module includes the enhanced recognition sets of each text in the text recognition range, and enters the usage stage. When the user inputs text again, the text is directly recognized locally through the KNN classification algorithm and then corrected.
[0092] In other embodiments, it is also possible to implement handwritten text recognition on other intelligent devices (such as smart wristbands, etc.) through the handwritten text recognition system and method provided in this application for human-computer interaction.
[0093] This embodiment takes the recognition of 26 capital English letters from "A" to "Z" as an example to illustrate the specific steps of the handwritten text recognition method proposed in this application. In other embodiments, it is also possible to recognize common texts such as handwritten lowercase English letters, numbers, Chinese characters, Japanese, Spanish, etc. through the handwritten text recognition method proposed in this application.
[0094] A handwritten text recognition method based on voice signals, as Figure 1 shown, includes the following steps:
[0095] S1: Set the text recognition range W. The text recognition range W includes M (M>0) different texts, W = {w1, w2,..., w M}; respectively generate the preliminary recognition sets of texts w1, w2,..., w M : I w1 、I w2 、...I wM , where I wi (i>1) stores: representing the text w iThe label that can be stably recognized, the characteristic sound signal, the classification data set corresponding to the characteristic sound signal one by one, and the preliminary recognition result of the characteristic sound signal is w i Initialize all M preliminary recognition sets as empty sets, where the labels are all initialized as not being able to be stably recognized.
[0096] In this embodiment, the text recognition range W is the 26 capital English letters from "A" to "Z", including 26 different letters. Generate the preliminary recognition set I of the character "A" respectively A and the preliminary recognition set I of the character "B" B and so on, until the preliminary recognition set I of the character "Z" z . All preliminary recognition sets are initialized as empty sets, and the label is set to 0, indicating that it cannot be stably recognized.
[0097] S2: Obtain a text picture set, which includes M handwritten text picture sets: H w1 、H w2 、…H wM , where the handwritten text picture in H wi (i > 1) is the handwritten text picture of the character w i . Each handwritten text picture set includes N (N > 2) handwritten text pictures of different users; convert each handwritten text picture in the M handwritten text picture sets into a sound signal respectively to obtain M sound signal sets: Y w1 、Y w2 、…Y wM , corresponding to the characters w1, w2,..., w M respectively; based on the M sound signal sets, construct a random forest model, the input of the random forest model is the sound signal, and the output is multiple model recognition results; enter the initialization stage.
[0098] Step S2 includes the following steps:
[0099] S21: Obtain multiple handwritten text pictures of a user, perform stretching and scaling operations on them to unify the specifications of the handwritten text pictures, and use the text in each handwritten text picture as the corresponding text of the handwritten text picture;
[0100] S22: Traverse all the handwritten text pictures of the user, put the handwritten text pictures with the corresponding text w1 into the handwritten text picture set H of the character w1 w1 、put the handwritten text pictures with the corresponding text w2 into the handwritten text picture set H of the character w2 w2 、…、put the handwritten text pictures with the corresponding text w M into the handwritten text picture set H of the character w M wM ;
[0101] S23: Repeat step S21 until each handwritten text picture set stores handwritten text pictures of N different users, obtaining a text picture set including M handwritten text picture sets;
[0102] S24: Select a handwritten text picture set from the text picture set, and convert each handwritten text picture in the handwritten text picture set into a sound signal through the particle swarm optimization algorithm, obtaining a sound signal set converted from the handwritten text picture set;
[0103] S25: Repeat step S24 until all handwritten text picture sets are converted into sound signal sets, obtaining M sound signal sets: Y w1 、Y w2 、…Y wM ,corresponding to characters w1, w2, …, w M ;
[0104] S26: Build a random forest model based on the M sound signal sets. The input of the random forest model is the sound signal, and the output is multiple model recognition results;
[0105] S27: Enter the initialization phase.
[0106] In this embodiment, the eMNIST dataset including a large number of handwritten letter pictures of different users is obtained from the Internet, and 26 handwritten text picture sets are established, N = 1000. In other embodiments, the size of N can be changed according to the accuracy requirements. Each handwritten text picture set includes handwritten text pictures from 1000 different users. For example, in the handwritten text picture set H A stores handwritten text pictures of capital letter 'A' from 1000 different users. In actual operation, it is possible that multiple handwritten text pictures in a handwritten text picture set come from the same user. For example, in the handwritten text picture set H A there are 1002 handwritten text pictures of capital letter 'A', among which 3 handwritten text pictures of capital letter 'A' come from the same user, and the remaining 999 handwritten text pictures of capital letter 'A' come from 999 different users respectively. For the sake of illustration, in this embodiment, each handwritten text picture set contains 1000 handwritten text pictures, and each handwritten text picture comes from a different user.
[0107] Step S24 includes the following steps:
[0108] S241: Select a handwritten text picture set from the text picture set;
[0109] S242: Obtain a handwritten text image from the handwritten text image set, perform a skeletonization operation on the handwritten text image and traverse it to obtain the range, starting point, ending point, and number of strokes of the text strokes in the handwritten text image;
[0110] S243: Generate a particle swarm. The initial position of each particle in the particle swarm is the starting point of the text stroke. According to the range and number of strokes of the text stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the velocity attributes of each particle, and increment the number of particle movements by 1;
[0111] S244: Update the positions of the particles. If the position of a particle is not within the range of the text stroke, remove the particle from the particle swarm. According to the range and number of strokes of the text stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the velocity attributes of each particle, and increment the number of particle movements by 1;
[0112] S245: If the number of movements is equal to the number of strokes, execute step S246; otherwise, repeat step S244;
[0113] S246: Obtain the movement trajectories of all the particles in the current particle swarm and perform an averaging operation to obtain the average trajectory, and calculate the time-distance sequence of the average trajectory as the sound signal corresponding to the handwritten text image;
[0114] S247: Repeat step S242 until all the handwritten text images in the handwritten text image set are converted into sound signals, obtaining the sound signal set converted from the handwritten text image set.
[0115] In this embodiment, through the conversion of the particle swarm algorithm, the sound signal corresponding to each handwritten text image is obtained, and then M sound signal sets are obtained. Each sound signal set contains 1000 sound signals. For example, the sound signal set Y A stores the sound signals converted from 1000 handwritten text images of the capital letter 'A'.
[0116] S3: Collect the original sound signal of the user during handwritten text input and preprocess it to obtain the characteristic sound signal. The characteristic sound signal is a time-distance sequence, specifically the distance between the user's hand and the sound signal receiving device at different times. If all the characters within the text recognition range W can be stably recognized, the initialization stage ends and the usage stage is entered, and step S7 is executed; otherwise, the initialization stage continues and step S4 is executed.
[0117] In step S2, the handwritten text image is converted into a sound signal, that is, a time-distance sequence, which can describe the distance relationship at any moment; in step S3, the position change of the user's hand relative to the sound signal receiving device during handwritten text input is collected to sense the user's text input action. In this way, in subsequent steps, through a certain method, the sound signal closest to the user's text input action can be found among the M sound signal sets, and the text corresponding to this sound signal can be considered as the text input by the user.
[0118] Step S3 includes the following steps:
[0119] S31: Send a high-frequency sound signal to the human hand through the sound signal generating device, and receive the sound signal reflected by the human hand through the sound signal receiving device and use it as the original sound signal;
[0120] S32: Perform normalization processing, Gaussian filtering, and noise reduction processing on the original sound signal;
[0121] S33: Obtain a time-distance sequence as the characteristic sound signal;
[0122] S34: If all the characters within the text recognition range W can be stably recognized, the initialization phase ends, enter the usage phase, and execute step S7; otherwise, continue the initialization phase and execute step S4.
[0123] When starting text recognition, it is still in the initialization phase, and there is no enhanced recognition set for each character in the text recognition range in the local storage module. Therefore, the obtained characteristic sound signal is input into the cloud computing module. After using it for a period of time, the local storage module stores the enhanced recognition sets for each character in the text recognition range. Enter the usage phase and directly use the KNN classification algorithm in the local enhanced recognition set for text recognition.
[0124] In the initialization phase, execute steps S4 - S6, perform text input and recognition through the random forest model and user confirmation, and at the same time obtain the enhanced recognition sets for different characters until all the characters within the text recognition range W can be stably recognized (that is, obtain the enhanced recognition sets for all characters within the text recognition range W), and the initialization phase ends.
[0125] S4: Input the characteristic sound signal into the random forest model to obtain multiple model recognition results, and the user confirms the preliminary recognition result m (m ∈ W) corresponding to the characteristic sound signal from the multiple model recognition results; obtain the preliminary recognition set I of character m m , and judge whether character m is stably recognized according to the labels in the preliminary recognition set I m . If so, output the preliminary recognition result m and execute step S3; if not, execute step S5.
[0126] Step S4 includes the following steps:
[0127] S41: Input the characteristic sound signal into the random forest model to obtain multiple model recognition results;
[0128] S42: The user finds the text entered in step S3 from the multiple model recognition results and confirms it as the preliminary recognition result m (m ∈ W) corresponding to the characteristic sound signal;
[0129] S43: Obtain the preliminary recognition set I m of the text m. According to the labels in the preliminary recognition set I m , determine whether the text m is stably recognized. If yes, output the preliminary recognition result m and execute step S3; if not, execute step S5.
[0130] S5: Write the characteristic sound signal into the preliminary recognition set I m ; Calculate the characteristic sound vector S m based on all the characteristic sound signals in the preliminary recognition set I user , obtain the sound signal set r m corresponding to the text m, and calculate the comprehensive similarity between the characteristic sound vector S user and all the sound signals in the sound signal set Y m respectively. Select the L (L > 1) sound signals with the largest comprehensive similarity in the sound signal set Y m as the latest classification data set of the text m, and write the latest classification data set into the preliminary recognition set I m and make it correspond to the characteristic sound signal. In this embodiment, L is taken as 500. In other embodiments, the size of L can be changed according to the size of N, the size of the local storage module, the accuracy requirement, the time limit, etc.
[0131] Step S5 includes the following steps:
[0132] S51: Write the characteristic sound signal into the preliminary recognition set I m ;
[0133] S52: Obtain all the characteristic sound signals in the preliminary recognition set I m , and calculate the characteristic sound vector S user :
[0134]
[0135] where S j represents the characteristic sound signal in the preliminary recognition set I m , n represents the number of characteristic sound signals in the preliminary recognition set I m ;
[0136] S53: Obtain the set of voice signals Y corresponding to the text m m , and calculate the characteristic voice vectors S respectively user and the Euclidean distances and Pearson similarities between S and all the voice signals in the set of voice signals Y m ;
[0137] S54: Select a voice signal p from the set of voice signals Y (the voice signal p ∈ Y m ), and calculate the comprehensive similarity Distance between the characteristic voice vector S m and the voice signal p: user
[0138]
[0139] wherein, S user is the characteristic voice vector, Image p is the voice signal p, d user_p is the Euclidean distance between the characteristic voice vector and the voice signal p, max d is the maximum value of the Euclidean distances between the characteristic voice vector and all the voice signals, min d is the minimum value of the Euclidean distances between the characteristic voice vector and all the voice signals, Pearson(S user , Image p ) is the Pearson similarity between the characteristic voice vector and the voice signal p, and α and β are weight coefficients;
[0140] In this embodiment, considering both the Euclidean distance and the Pearson similarity, the weight coefficients α and β are both taken as 0.5, which can evaluate the similarity more objectively.
[0141] S55: Repeat step S54 until the comprehensive similarities between the characteristic voice vector S user and all the voice signals in the set of voice signals Y m are obtained;
[0142] S56: Select the L voice signals with the largest comprehensive similarities in the set of voice signals Y m as the latest classification data set of the text m, and write the latest classification data set into the preliminary recognition set I m and make it correspond to the characteristic voice signal.
[0143] S6: Determine whether the preliminary recognition set I m reaches a stable state. If it reaches a stable state, then the preliminary recognition set I m The label is updated to be stably recognizable, that is, the text m can be stably recognized. The latest classification dataset is used as the enhanced recognition set for the text m, and the preliminary recognition result m is output. Then, step S3 is executed; if the stable state is not reached, the preliminary recognition result m is directly output, and step S3 is executed.
[0144] Assume that the user's input text is "A", and it is the user's first input of "A". Obviously, there is no enhanced recognition set for the text "A" at this time. First, multiple model recognition results are obtained through the random forest model. The model recognition results are "A", "H", "N". The user finds their input "A" in the model recognition results and confirms it. "A" is the preliminary recognition result and is directly output. At the same time, the preliminary recognition set of "A" is an empty set at this time. First, the characteristic sound signal is written into the preliminary recognition set, and then the latest classification dataset is obtained and written into the preliminary recognition set. After that, step S6 is executed to determine whether the preliminary recognition set reaches the stable state.
[0145] Assume that the user's input text is "D", and it is the user's tenth input of "D". At this time, only the enhanced recognition sets for the texts "A", "C", "Y" are obtained. Therefore, it is still in the initialization stage. First, multiple model recognition results are obtained through the random forest model. The model recognition results are "O", "Q", "D". The user finds their input "D" in the model recognition results and confirms it. "D" is the preliminary recognition result and is directly output. At the same time, the preliminary recognition set of "D" stores the characteristic sound signals and classification datasets of the previous nine inputs of the text "D", and the label is 0, indicating that it cannot be stably recognized. First, the characteristic sound signal is written into the preliminary recognition set, and then the latest classification dataset is obtained and written into the preliminary recognition set. After that, step S6 is executed to determine whether the preliminary recognition set reaches the stable state.
[0146] Assume that the user's input text is "U", and it is the user's twelfth input of "U". At this time, the enhanced recognition set for the text "U" has been obtained, but the enhanced recognition sets for the texts "P", "R" have not been obtained yet, and it is still in the initialization state. First, multiple model recognition results are obtained through the random forest model. The model recognition results are "U", "V", "Y". The user finds their input "U" in the model recognition results and confirms it. "U" is the preliminary recognition result and is directly output. However, since the enhanced recognition set for the text "U" has been obtained, step S3 is directly executed to wait for the user's next input.
[0147] When calculating the comprehensive similarity, first obtain all the characteristic voice signals in the preliminary recognition set, that is, all the characteristic voice signals input by the user before, calculate the average value as the characteristic voice vector, and then calculate the comprehensive similarity between the characteristic voice vector and the voices in the voice signal set. This can synthesize historical data and better reflect the user's handwritten text habits. The higher the comprehensive similarity, the closer it is to the user's handwriting habits. In this embodiment, L is taken as 500, that is, the 500 voice signals with the highest comprehensive similarity are selected as the latest classification data set.
[0148] Step S6 includes the following steps:
[0149] S61: Obtain all the classification data sets in the preliminary recognition set I m If the number of classification data sets is equal to the preset recognition threshold F1 (F1>2), then the preliminary recognition set I m reaches a stable state, and step S64 is executed; if the number of classification data sets is greater than or equal to the preset judgment threshold Kp (F1>Kp≥2) and less than the preset recognition threshold F1, then step S62 is executed; if the number of classification data sets is less than the preset judgment threshold Kp, then step S65 is executed;
[0150] S62: Obtain the latest Kp classification data sets written in the preliminary recognition set I m ;
[0151] S63: Calculate the change error between the Kp classification data sets. If the change error is less than the preset stability threshold, then the preliminary recognition set I m reaches a stable state, and step S64 is executed. Otherwise, the preliminary recognition set I m has not reached a stable state, and step S65 is executed;
[0152] S64: Update the label of the preliminary recognition set I m to be stably recognized, that is, the text m can be stably recognized. Take the latest classification data set as the enhanced recognition set of the text m, output the preliminary recognition result m, and execute step S3;
[0153] S65: Output the preliminary recognition result m, and execute step S3.
[0154] In this embodiment, F1 is taken as 10 and Kp is taken as 2. There are two criteria for the preliminary recognition set to reach a stable state. One is that there are already 10 characteristic voice signals in the preliminary recognition set, indicating that the user has input the text 10 times. At this time, the latest written classification data set is obtained by calculating the comprehensive similarity based on the average value of the previous 10 characteristic voice signals, which can represent the user's handwriting habit. When the preliminary recognition set reaches a stable state, the latest written classification data set is used as the enhanced recognition set for this text. The other is that the change error of the latest written 2 classification data sets is very small, such as only 1 or 2 voice signals being different. In this way, it is also considered that the preliminary recognition set reaches a stable state, and the latest written classification data set is used as the enhanced recognition set for this text.
[0155] In other embodiments, the values of F1 and Kp can also be changed according to the accuracy requirements and actual usage conditions.
[0156] When the preliminary recognition set reaches a stable state, that is, the text can be stably recognized, the enhanced recognition set is written into the local storage module.
[0157] In this way, when M texts within the text recognition range W are all stably recognized, the local storage module contains the enhanced recognition sets of these M texts. The initialization phase ends and enters the usage phase. After the user inputs the text again, the voice signal closest to the characteristic voice signal in the M enhanced recognition sets is directly obtained locally through the KNN classification algorithm, so as to perform text recognition.
[0158] S7: Obtain the enhanced recognition sets of the texts w1, w2, …, w M and use the voice signal closest to the characteristic voice signal in the M enhanced recognition sets as the similar voice signal based on the KNN classification algorithm, and take the corresponding text of the similar voice signal as the recognition result;
[0159] S8: Repeat step S3 until the recognition of the handwritten text ends, and perform error correction processing on all the obtained recognition results to obtain the final recognized text.
[0160] In step S8, the text error correction data packet based on natural language processing is used to perform error correction on all the obtained recognition results at the word dimension to obtain the final recognized text. In this embodiment, the text error correction data packet based on natural language processing provided by the Android mobile phone is used to perform error correction on the words that may be misclassified by the user at the word dimension, making the finally recognized text more accurate.
[0161] It further includes step S9, and step S9 is specifically: obtain the feedback result of the user on the recognition accuracy of the handwritten text; if the feedback result is that the accuracy rate is lower than the pre-set accuracy rate threshold, then execute step S1, otherwise, execute step S3.
[0162] After long-term use, if the recognition accuracy is low, or the user's handwriting habit changes, step S1 can be executed again to improve the user experience.
[0163] To verify the accuracy and reliability of this application, a test experiment was conducted. The experiment content was that 15 users used this application to write texts in different languages at different locations, and the system collected a large amount of data. By comparing the classification of the system's real-time results with the users' real texts, the results are as Figure 3 and Figure 4 shown. Figure 3 The accuracy rates of 15 users writing handwritten numbers, capital English, lowercase English, Chinese, and Japanese are as Figure 3 shown. For this application, the recognition accuracy of texts in any language written by users with any handwriting habit can reach over 93%. Among them, the recognition accuracy of lowercase English can reach 97%, and the recognition accuracy of Chinese can reach 94.2%. Figure 4 It is a graph of the accuracy rates of 15 users using this application in different scenarios. It can be seen from the graph that in a quiet environment, this application can achieve very good results in both single-character recognition and word recognition, with an accuracy rate above 95%. In a relatively noisy environment such as a station or subway, the single-character recognition effect of the system will decline, but after the error correction by the error correction unit of this application, the word accuracy rate of the system can still reach above 95%.
[0164] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of this application based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the existing technology should be within the protection scope determined by the claims.
Claims
1. A handwritten text recognition method based on sound signals, characterized in that Including the following steps: S1: Set the text recognition range W, where the text recognition range W includes M, M > 0 different characters, W = {w1, w2, …, w M}; Generate preliminary recognition sets for characters w1, w2, …, w M respectively: I w1 , I w2 , … I wM , where I wi , i > 1 stores: a label indicating whether the character w i is stably recognized, a characteristic sound signal, and a classification data set corresponding one-to-one with the characteristic sound signal. The preliminary recognition result of the characteristic sound signal is w i ; Initialize all M preliminary recognition sets to empty sets, where the labels are all initialized as not stably recognized; S2: Obtain a text image set, where the text image set includes M handwritten text image sets: H w1 , H w2 , … H wM , where H wi , i > 1, is a handwritten text image of the character w i . Each handwritten text image set includes N handwritten text images of different users, where N > 2. Convert each handwritten text image in the M handwritten text image sets into a sound signal respectively to obtain M sound signal sets: Y w1 , Y w2 , … Y wM , corresponding to the characters w1, w2, …, w M respectively. Based on the M sound signal sets, construct a random forest model, where the input of the random forest model is a sound signal and the output is multiple model recognition results. Enter the initialization stage; S3: Collect the original voice signal of the user during handwritten text input, and preprocess it to obtain a characteristic voice signal, where the characteristic voice signal is a time-distance sequence; if all the characters within the text recognition range W can be stably recognized, the initialization phase ends, enter the usage phase, and execute step S7, otherwise, continue the initialization phase and execute step S4; S4: Input the characteristic sound signal into the random forest model to obtain multiple model recognition results. The user confirms the preliminary recognition result m corresponding to the characteristic sound signal from the multiple model recognition results, where m ∈ W; obtain the preliminary recognition set I of the text m m , and judge whether the text m is stably recognized according to the labels in the preliminary recognition set I m . If yes, output the preliminary recognition result m and execute step S3; if not, execute step S5; S5: Write the characteristic sound signal into the preliminary recognition set I m ; Based on the preliminary identification set I m All characteristic sound signals in the calculation result are used to obtain the characteristic sound vector S user , get the sound signal set Y corresponding to the text m m , respectively calculate the characteristic sound vector S user With sound signal set Y m The comprehensive similarity between all sound signals in , select the sound signal set Y m The L sound signals with the largest comprehensive similarity among the L sound signals, L>1, are used as the latest classification data set of the text m, and the latest classification data set is written into the preliminary recognition set I m and corresponds to the characteristic sound signal; S6: Determine whether the initial recognition set I m has reached a stable state. If it has reached a stable state, update the label of the initial recognition set I m to be stably recognized, that is, the character m can be stably recognized. Use the latest classified data set as the enhanced recognition set of the character m, output the initial recognition result m, and execute step S3; if it has not reached a stable state, directly output the initial recognition result m and execute step S3; S7: Obtain the enhanced recognition sets of words w1, w2, …, w M , and based on the KNN classification algorithm, obtain the voice signals in the M enhanced recognition sets that are closest to the feature voice signal as similar voice signals, and use the corresponding words of the similar voice signals as the recognition results; S8: Repeat step S3 until the handwritten text recognition ends, perform error correction processing on all the obtained recognition results and obtain the final recognized text; The step S5 includes the following steps: S51: Write the characteristic sound signal into the preliminary recognition set I m ; S52: Obtain the initial recognition set I m All the characteristic sound signals in it, and calculate the characteristic sound vector S user : Among them, S j represents the characteristic voice signal in the preliminary recognition set I m , n represents the number of characteristic voice signals in the preliminary recognition set I m ; S53: Obtain the set of voice signals Y corresponding to the text m m , and calculate the characteristic voice vectors S respectively user and the Euclidean distance and Pearson similarity between all voice signals in the set of voice signals Y m ; S54: Select a voice signal p from the voice signal set Y m where the voice signal p ∈ Y m and calculate the comprehensive similarity Distance between the characteristic voice vector S user and the voice signal p: Among them, S user is the characteristic sound vector, Image p is the sound signal p, d user_p is the Euclidean distance between the characteristic sound vector and the sound signal p, max d is the maximum value of the Euclidean distances between the characteristic sound vector and all sound signals, min d is the minimum value of the Euclidean distances between the characteristic sound vector and all sound signals, Pearson(S user , Image p ) is the Pearson similarity between the characteristic sound vector and the sound signal p, and α and β are weight coefficients; S55: Repeat step S54 until the characteristic sound vector S is obtained user and the sound signal set Y m the comprehensive similarity between all sound signals S56: Select the voice signal set Y m The L voice signals with the largest comprehensive similarity in are used as the latest classification data set of the text m, and the latest classification data set is written into the preliminary recognition set I m and corresponds to the characteristic voice signal.
2. The handwritten text recognition method based on sound signals according to claim 1, wherein The step S2 includes the following steps: S21: Obtain multiple handwritten text pictures of a user, perform stretching and scaling operations on them to unify the specifications of the handwritten text pictures, and use the characters in each handwritten text picture as the corresponding characters of the handwritten text picture; S22: Traverse all the handwritten text images of the user, and put the handwritten text images with the corresponding text w1 into the handwritten text image set H of the text w1 w1 , put the handwritten text images with the corresponding text w2 into the handwritten text image set H of the text w2 w2 , …, put the handwritten text images with the corresponding text w M into the handwritten text image set H of the text w M ; wM ; S23: Repeat step S21 until each handwritten text picture set stores handwritten text pictures of N different users respectively, and obtain a text picture set including M handwritten text picture sets; S24: Select a handwritten text picture set from the text picture set, and convert each handwritten text picture in the handwritten text picture set into a voice signal through a particle swarm algorithm to obtain a voice signal set converted from the handwritten text picture set; S25: Repeat step S24 until all the handwritten text picture sets are converted into sound signal sets, obtaining M sound signal sets: Y w1 、Y w2 、…Y wM ,respectively corresponding to the characters w1, w2, …, w M ; S26: Build a random forest model based on the M voice signal sets, where the input of the random forest model is the voice signal and the output is multiple model recognition results; S27: Enter the initialization phase.
3. The handwritten text recognition method based on voice signals according to claim 2, characterized in that The step S24 includes the following steps: S241: Select a handwritten text picture set from the text picture set; S242: Obtain a handwritten text picture from the handwritten text picture set, perform skeletonization operation on the handwritten text picture and traverse it to obtain the range, starting point, ending point and number of strokes of the characters in the handwritten text picture; S243: Generate a particle swarm, the initial position of each particle in the particle swarm is the starting point of the character stroke, according to the range and number of strokes of the character stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the speed attributes of each particle, and the number of particle movements is incremented by 1; S244: Update the position of the particle. If the position of a particle is not within the range of the character stroke, remove the particle from the particle swarm, according to the range and number of strokes of the character stroke, obtain the direction range of the next stage, randomly generate the direction attributes of each particle within the direction range, randomly generate the speed attributes of each particle, and the number of particle movements is incremented by 1; S245: If the number of times is equal to the number of strokes, execute step S246, otherwise, repeat step S244; S246: Obtain the movement trajectories of all the particles in the current particle swarm and perform an averaging operation to obtain an average trajectory, and calculate the time-distance sequence of the average trajectory as the voice signal corresponding to the handwritten text picture; S247: Repeat step S242 until all the handwritten text pictures in the handwritten text picture set are converted into voice signals to obtain a voice signal set converted from the handwritten text picture set.
4. A handwritten text recognition method based on sound signals according to claim 1, characterized in that, The step S3 includes the following steps: S31: Send a high-frequency sound signal to the human hand through a sound signal generating device, and receive the sound signal reflected by the human hand through a sound signal receiving device and use it as the original sound signal; S32: Perform normalization processing, Gaussian filtering, and noise reduction processing on the original sound signal; S33: Obtain a time-distance sequence as the characteristic sound signal; S34: If all the characters within the text recognition range W can be stably recognized, the initialization stage ends and enters the usage stage, and step S7 is executed; otherwise, the initialization stage continues and step S4 is executed.
5. A method for handwritten text recognition based on sound signals according to claim 1, wherein The step S4 includes the following steps: S41: Input the characteristic sound signal into a random forest model to obtain multiple model recognition results; S42: The user finds and confirms the text input in step S3 among the multiple model recognition results as the preliminary recognition result m corresponding to the characteristic sound signal, where m ∈ W; S43: Obtain the preliminary recognition set I of the text m m , and determine whether the text m is stably recognized according to the labels in the preliminary recognition set I m . If yes, output the preliminary recognition result m and execute step S3; if no, execute step S5.
6. A handwritten text recognition method based on sound signals according to claim 1, characterized in that The step S6 includes the following steps: S61: Obtain the initial recognition set I m for all classification data sets in, if the number of classification data sets is equal to the preset recognition threshold F1, where F1 > 2, then the initial recognition set I m reaches a stable state, and step S64 is executed; if the number of classification data sets is greater than or equal to the preset judgment threshold Kp, where F1 > Kp ≥ 2 and less than the preset recognition threshold F1, then step S62 is executed; if the number of classification data sets is less than the preset judgment threshold Kp, then step S65 is executed; S62: Obtain the initial recognition set I m The Kp most recently written classification data sets; S63: Calculate the change error between Kp classification data sets. If the change error is less than a preset stable threshold, then the preliminary identification set I m reaches a stable state, and step S64 is executed. Otherwise, the preliminary identification set I m does not reach a stable state, and step S65 is executed; S64: Update the labels in the preliminary recognition set I m to those that can be stably recognized, i.e., the text m can be stably recognized. Use the latest classified data set as the enhanced recognition set of the text m, output the preliminary recognition result m, and execute step S3; S65: Output the preliminary recognition result m and execute step S3.
7. A method for handwritten text recognition based on voice signals according to claim 1, characterized in that, In the step S8, use a word-level error correction data packet based on natural language processing to correct the obtained recognition results to obtain the final recognized text.
8. A method for handwritten text recognition based on sound signals according to claim 1, characterized in that, It further includes step S9, and the step S9 is specifically: obtain the feedback result of the user on the recognition accuracy of the handwritten text; if the feedback result is that the accuracy is lower than the preset accuracy threshold, execute step S1; otherwise, execute step S3.
9. A handwritten text input recognition system based on sound signals, characterized in that, Based on the handwritten text recognition method described in any one of claims 1-8, it includes a control module, a local sound signal acquisition module, a local storage module, and a cloud computing module; The control module, the local sound signal acquisition module, and the local storage module are integrated on the user side, and the cloud computing module is set in the cloud; The control module is respectively connected to the local sound signal acquisition module and the local storage module, and the control module is communicatively connected to the cloud computing module; The local sound signal acquisition module includes a sound signal generating device, a sound signal receiving device, and a data preprocessing unit. The sound signal generating device is used to send a high-frequency sound signal to the human hand, and the sound signal receiving device is used to receive the sound signal reflected by the human hand and transmit it as the original sound signal to the data preprocessing unit; The data preprocessing unit is used to preprocess the original sound signal to obtain the characteristic sound signal; The local storage module is used to store the enhanced recognition set; The cloud computing module includes a data calculation unit and a storage unit connected to each other; the data calculation unit is used for data calculation during the text recognition process; the storage unit is used to store the text picture set and the data during the text recognition process.
Citation Information
Patent Citations
Handwritten text recognition methods, systems, devices and media
CN110866499B