Handwritten character recognition method and system based on convolutional neural network

By analyzing the performance of speech data and character input, intelligently selecting the data input method, the problem of poor recognition effect when combining speech input and character input in the prior art is solved, and more efficient and accurate handwritten character recognition is achieved.

CN120182985AInactive Publication Date: 2025-06-20JIANGSU VOCATIONAL COLLEGE OF BUSINESS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296312.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing handwritten character recognition method based on convolutional neural networks is mainly aimed at static images, ignoring that users may obtain learning materials through speech input combined with character input, resulting in the accuracy of speech input and the convenience of character input affecting the recognition effect.

Method used

By collecting noise and device temperature information in voice data, analyzing the accuracy of voice input, and selecting different data input methods based on the accuracy. At the same time, the speed and error rate of character input are detected, the accuracy score and contrast are analyzed, and the accuracy of character input is comprehensively evaluated.

Benefits of technology

It improves the accuracy and efficiency of handwritten character recognition, and ensures the accuracy of acquisition of learning materials, especially optimization in different environments and user input habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182985A_ABST
    Figure CN120182985A_ABST
Patent Text Reader

Abstract

The invention discloses and relates to the technical field of computer image processing, in particular to a handwritten character recognition method and a handwritten character recognition system which can acquire learning materials by combining voice input with character input and are based on a convolutional neural network, and the method comprises the following steps: acquiring voice data, extracting and quantifying noise through a sound pressure level, and analyzing noise spectrum change. Analyzing voice input accuracy according to the sound pressure level and the noise spectrum change, selecting a voice input mode or adding character input verification on the basis of the voice input mode, detecting a character input accuracy score, obtaining a contrast ratio, comprehensively analyzing the accuracy score and the contrast ratio, and evaluating character input to obtain an accuracy degree percentage of learning materials; according to the noise data, the equipment temperature and the voice recognition performance obtained in real time, the accuracy is predicted through the support vector machine, when the prediction result shows that the accuracy is low, the user is automatically reminded to perform character input verification, and the accuracy of the finally obtained learning material is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer image processing, and more specifically, to a handwritten character recognition method and system based on a convolutional neural network. Background Art

[0002] With the rapid development of artificial intelligence technology, significant progress has been made in handwritten character recognition methods. Traditional handwritten character recognition methods mainly rely on pattern matching and feature extraction, but these methods have certain limitations when dealing with complex backgrounds, different writing styles, and noise interference. In recent years, as an important technology in the field of deep learning, convolutional neural networks have shown great potential in the field of handwritten character recognition due to their powerful feature extraction and pattern recognition capabilities.

[0003] However, most of the existing handwritten character recognition methods based on convolutional neural networks are for static image recognition, ignoring the fact that in actual application scenarios, users may obtain learning materials by combining voice input with character input. In this case, the accuracy of voice input and the convenience of character input become key factors affecting the handwritten character recognition effect.

[0004] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a handwritten character recognition method and system based on a convolutional neural network to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A handwritten character recognition method and system based on a convolutional neural network, including the following steps: S1, collect voice data, extract noise in the voice data, and perform quantization processing on the noise through sound pressure level; collect the temperature of the recording device when collecting the voice data, and the number of noise types extracted, and obtain the spectral change of the noise according to the temperature of the recording device and the number of noise types; S2, analyze the accuracy of obtaining learning materials through voice input by sound pressure level and spectral change of the noise, select different data input methods according to the high or low accuracy, adopt the voice input method for high accuracy, and add character input verification on the basis of the voice input method for low accuracy; S3, detect the speed and error rate during character input, and analyze the accuracy score of character input; obtain the contrast of character input by comparing the text of character input with the text converted from voice data; S4, comprehensively analyze the accuracy score and contrast to evaluate the percentage of the accuracy of character input for obtaining learning materials.

[0008] In a preferred embodiment, the temperature of the recording device when acquiring voice data is real-time collected through the integrated temperature sensor of the recording device. By converting the voice data into two-dimensional spectrogram features, the convolutional neural network (CNN) can capture the noise features in the voice data and perform noise classification based on this, and count the number of noise types according to the captured noise features in the voice data; the voice data is processed by equal-time segmentation to obtain multiple small audio segments, the noise of each small audio segment is quantized in sound pressure level to obtain the noise sound pressure level, the noise of each small audio segment is spectro-processed to obtain the maximum spectral point and the minimum spectral point, and the difference between the maximum spectral point and the minimum spectral point is obtained to get the spectral change value of the noise of each small audio segment; wherein, the calculation formula for the noise sound pressure level of each small audio segment is: = 20· , where is the sound pressure level; is the sound pressure at the measurement point; is the reference sound pressure; the calculation formula for the spectral change value of the noise of each small audio segment is: ΔX = , where ΔX is the spectral change value; the maximum spectral point = , is the frequency corresponding to the maximum amplitude value, is the amplitude at this frequency point, the minimum spectral point = , is the frequency corresponding to the minimum amplitude value, is the amplitude value at this frequency point, and are the maximum and minimum amplitude values respectively.

[0009] Specifically, the real-time collection of temperature data by using the integrated temperature sensor of the recording device can more accurately reflect the environmental conditions when the voice data is picked up, and thus perform a more refined classification and prediction of the noise features. This helps to improve the stability and recognition accuracy of the system in different temperature environments

[0010] In a preferred embodiment, the different noise sound pressure levels obtained from multiple small audio segments are combined into a sound pressure level data set, and the different spectral change values obtained from multiple small audio segments are combined into a spectral change data set. When acquiring learning materials, the support vector machine model in the linear regression model is used to predict the accuracy of the learning materials obtained through voice input. The sound pressure level data set and the spectral change data set are used as input variables to construct a support vector machine model, and the support vector machine model constructed by the sound pressure level data set and the spectral change data set is used to predict the accuracy of the learning materials obtained through voice input.

[0011] Specifically, by constructing a support vector machine model and using the sound pressure level dataset and the spectrum change dataset as input variables, the accuracy of the voice input method is predicted. This method can intelligently evaluate the quality of voice input, provide a basis for the subsequent selection of data input methods, and help improve the overall recognition efficiency and accuracy.

[0012] In a preferred embodiment, the accuracy score of the learner's character input is obtained through comprehensive analysis of the speed and error rate during character input. The speed during character input is measured by calculating the number of characters input per unit time. The formula for the speed S of character input is: S = ; The error rate The calculation formula is: = ; Then, the calculation formula for the correct rate W is: W = 1 - ; Where S is the number of characters with substitution errors, that is, the number of different characters between the characters input by the learner and the correct character output; D is the number of characters with deletion errors, that is, the number of characters missed by the learner's character input; I is the number of characters with insertion errors, that is, the number of extra characters in the learner's character input; N is the total number of characters in the reference text, that is, the number of characters in the correct text. The value of the error rate is usually between 0 and 1. The smaller the value, the lower the error rate of the learner's character input.

[0013] Specifically, by calculating the speed and error rate of character input, the input performance of the learner can be more comprehensively reflected, providing more accurate data support for the subsequent evaluation of the accuracy of learning material acquisition.

[0014] In a preferred embodiment, the accuracy score of the learner's character input is represented by a comprehensive scoring formula, combining the speed S and the error rate E into a weighted scoring model: Accuracy score : = ; Where S is the speed, E is the error rate, (1 - E) is the correct rate W, , are the weights of the speed and the correct rate respectively; The words in the text of character input are combined into an input set, and the words in the text after converting the voice data are combined into a conversion set. The contrast of character input is obtained by calculating the ratio of the intersection and union of the two sets. The contrast formula is: J = ; Where is the input set, and B is the conversion set; is the size of the intersection of the two sets, indicating the number of words that appear together, is the size of the union of the two sets, indicating the total number of all different words, and J is the contrast.

[0015] Specifically, by combining speed and accuracy into a weighted scoring model, the accuracy of character input can be evaluated more comprehensively. This method takes into account both input speed and input accuracy, which helps to improve the accuracy and reliability of the evaluation results.

[0016] In a preferred embodiment, the percentage G of accuracy is calculated through a weighted formula. The specific formula is: Percentage G of accuracy: G = ( ) × 100%; where the accuracy score : = ; is the contrast J : J = ; is the weight of the accuracy score, is the weight of the contrast.

[0017] Specifically, through the weighted formula, a comprehensive analysis is performed on the accuracy score and the contrast to calculate the percentage of the accuracy of the learning materials obtained. This method can more comprehensively reflect the impact of character input on the acquisition of learning materials, providing more accurate feedback and guidance for learners.

[0018] A handwritten character recognition system based on a convolutional neural network is used to implement the handwritten character recognition method based on a convolutional neural network according to any one of claims 1-6. It is characterized in that it includes a data acquisition module, an input selection module, a character analysis module, and a comprehensive evaluation module; the data acquisition module is used to collect voice data, extract and quantify noise through sound pressure level, record the device temperature, extract the number of noise types, and analyze the change of the noise spectrum; the input selection module is used to analyze the accuracy of voice input according to the sound pressure level and the change of the noise spectrum, and select the voice input method or add character input verification on the basis of the voice input method; the character analysis module is used to detect the character input speed, error rate, analyze the accuracy score, and obtain the contrast by comparing the character input with the voice conversion text; the comprehensive evaluation module is used to comprehensively analyze the accuracy score and the contrast, and evaluate the percentage of the accuracy of the learning materials obtained by the character input.

[0019] The technical effects and advantages of the handwritten character recognition method and system based on a convolutional neural network of the present invention: This method analyzes the accuracy of voice input by collecting information such as noise and device temperature in voice data, and selects different data input methods according to the level of accuracy. At the same time, this method also detects the speed and error rate of character input, analyzes the accuracy score and contrast of character input, so as to comprehensively evaluate the accuracy of character input for obtaining learning materials.

[0020] The present invention predicts and optimizes the accuracy through a support vector machine based on the real-time acquired noise data, device temperature, and performance of speech recognition. When the prediction result shows low accuracy, the system can automatically remind the user to perform character input verification to ensure the accuracy of the finally obtained learning materials.

[0021] The present invention improves the accuracy and efficiency of handwritten character recognition by selecting the most suitable data input method and comprehensively evaluating according to the user's input performance. Brief Description of the Drawings

[0022] Figure 1 It is a schematic flow diagram of the handwritten character recognition method based on convolutional neural network of the present invention;

[0023] Figure 2 It is a schematic structural diagram of the handwritten character recognition system based on convolutional neural network of the present invention. Detailed Embodiments

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] Embodiment 1: The present invention discloses a handwritten character recognition method based on convolutional neural network, including the following steps:

[0026] S1. Collect speech data, extract the noise in the speech data, and perform quantization processing on the noise through the sound pressure level; collect the temperature of the recording device when collecting the speech data, and the number of noise types extracted, and obtain the spectral change of the noise according to the temperature of the recording device and the number of noise types;

[0027] S2. Analyze the accuracy of the learning materials obtained through the speech input method through the sound pressure level and the spectral change of the noise, select different data input methods according to the high or low accuracy, adopt the speech input method for high accuracy, and add character input verification on the basis of the speech input method for low accuracy;

[0028] S3. Detect the speed and error rate during character input, and analyze the accuracy score of character input; obtain the contrast of character input by comparing the text of character input with the text after conversion of speech data;

[0029] S4. Comprehensively analyze the accuracy score and contrast to evaluate the percentage of the accuracy of character input for obtaining learning materials.

[0030] In this embodiment, S1 includes the following processes:

[0031] Collect the learner's human voice recorded by the recording device and the background noise to form voice data. Use a noise cancellation algorithm to extract the noise in the voice data, and quantify the noise through the sound pressure level. The sound pressure level represents the intensity of the sound, and its formula is: = 20· , where is the sound pressure level (unit: dB); is the sound pressure at the measurement point; is the reference sound pressure, usually 20 micropascals, which is the lowest sound pressure that the human ear can hear. The measurement of the sound pressure level is usually carried out at a specific frequency to describe the strength of the sound. When the sound pressure level of the noise approaches or exceeds the voice signal of the learner's human voice, the signal-to-noise ratio becomes low, resulting in difficulty for the voice assistant to distinguish the voice signal. It should be noted that the voice assistant (VoiceAssistant) is an application program developed based on artificial intelligence technology. Through speech recognition, natural language processing, and speech synthesis technologies, it can interact with users through voice and provide services.

[0032] Through the integrated temperature sensor of the recording device, the temperature of the recording device when acquiring voice data is collected in real time. The working temperature of the recording device will affect the pickup effect of the voice data; by converting the voice data into a two-dimensional spectrogram or MFCC (Mel Frequency Cepstral Coefficients) features, the CNN (Convolutional Neural Network) can capture the noise features in the voice data and classify the noise on this basis. The CNN can classify and predict the noise type by learning different types of noise features, such as wind noise, white noise, background conversation, etc., and count the number of noise types according to the captured noise features in the voice data.

[0033] Perform equal-time segmentation processing on the voice data to obtain multiple small audio segments. Quantify the sound pressure level of the noise in each small audio segment to obtain the noise sound pressure level. Perform spectral processing on the noise in each small audio segment to obtain the maximum spectral point and the minimum spectral point, and calculate the difference between the maximum spectral point and the minimum spectral point to obtain the spectral change value of the noise in each small audio segment.

[0034] The total noise spectrum of each small audio segment is represented as the spectral superposition of each noise feature, and the spectrum of each noise feature is affected by the device temperature. The total noise spectrum is modeled as: = ; where N represents the number of noise types; is the spectrum of the th type of noise; is the temperature-dependent correction factor of the spectrum of the th type of noise with respect to temperature, obtained through the historical experimental database; To be specific , it can be assumed that the dependence of certain noise types on temperature is linear, especially the influence of thermal noise. Therefore, it can be expressed as: = 1 + ; where is the reference temperature (such as room temperature, 25 °C or 298 K); is the noise type 's sensitivity to temperature changes.

[0035] The maximum spectral point is obtained according to the frequency point corresponding to the maximum amplitude value in the spectrogram of the total noise spectrum. Among them, the maximum spectral point = , is the frequency corresponding to the maximum amplitude value, is the amplitude at this frequency point; the minimum spectral point is obtained according to the frequency point corresponding to the minimum amplitude value in the spectrogram of the total noise spectrum. Among them, the minimum spectral point = , is the frequency corresponding to the minimum amplitude value, is the amplitude value at this frequency point; therefore, the spectral change value is the difference between the maximum amplitude value and the minimum amplitude value, that is, ΔX = , where ΔX is the spectral change value; and are the maximum and minimum amplitude values respectively.

[0036] In this embodiment, S2 includes the following processes:

[0037] The system combines the different noise sound pressure levels obtained from multiple small audio segments into a sound pressure level data set, and combines the different spectral change values obtained from multiple small audio segments into a spectral change data set. When obtaining learning materials, the system uses the support vector machine model in the linear regression model to predict the accuracy of the learning materials obtained through voice input. The system constructs a support vector machine model with the sound pressure level data set and the spectral change data set as input variables, and predicts the accuracy of the learning materials obtained through voice input by constructing a support vector machine model with the sound pressure level data set and the spectral change data set. The specific steps are as follows:

[0038] Step A1: Normalize the data in the sound pressure level data set and the spectral change data set and use them as input features.

[0039] Step A2: Select the radial basis function as the kernel function to perform feature transformation on the input features and combine them into an array. The feature transformation formula can be: A(y,x)= Where A(y, x) is the calculation result after the support vector machine trains on the average values of the two normalized data sets. y is the average value of the normalized sound pressure level data in the sound pressure level data set, x is the average value of the normalized spectral change value data in the spectral change data set, and e is the natural logarithm base. is a hyperparameter.

[0040] It should be noted that when the support vector machine uses the radial basis kernel, the hyperparameter directly affects the shape of the decision boundary in the feature space and determines how the support vector machine "stretches" or "compresses" the similarity between data points. In this system, is used to adjust the sensitivity and fitting ability of the model to the data.

[0041] Step A3: Divide the training set and the test set. Randomly divide the sound pressure level data set and the spectral change data set into training sets and test sets with different proportions. Randomly select a set of data in the training set and the test set to perform calculations through the kernel function to obtain the training set calculation result and the test set calculation result.

[0042] Step A4: Use the loss function to converge the calculation results. The loss function formula can be: B(a, b) = max(0, 1 - ab), where B(a, b) is the loss amount of the kernel function, a is the training set calculation result, b is the test set calculation result, and converge the calculation results of the support vector machine according to the loss amount of the kernel function. The convergence formula is: (y, x) = (y, x) - B(a, b). By randomly selecting data in the training set and the test set multiple times to calculate the loss amount B(a, b) of the kernel function, the final convergence result (y, x) is obtained.

[0043] Step A5: Take the final convergence result (y, x) as the accuracy of the learning materials obtained through the voice input method.

[0044] It should be noted that the convergence result (y, x) represents the adaptation and fitting degree of the support vector machine model to the input data. If the convergence result (y, x) is large, it means that the model can better fit the data in the training set and the test set during the training process and can more accurately capture the relationship between the input features and the output results. By setting the threshold T, when (y, x) is greater than the threshold T, the system will consider that the accuracy of the learning materials obtained through voice input is high, and the prediction result is reliably of high accuracy. For high accuracy, the system directly uses the learner's voice data to obtain learning materials. When When (y, x) is less than the threshold T, the system considers that the accuracy of the learning materials obtained through voice input is low, and the prediction result is unreliable as low accuracy. While directly adopting the learner's voice data, the system reminds the learner to increase character input verification to obtain learning materials.

[0045] The system can predict and optimize the accuracy through a support vector machine based on the real-time acquired noise data, device temperature, and speech recognition performance. When the prediction result shows low accuracy, the system can automatically remind the user to perform character input verification to ensure the accuracy of the finally obtained learning materials. This adaptive accuracy evaluation method enables the system to optimize for different environments and different user input habits, improving the overall learning effect.

[0046] In this embodiment, S3 includes the following processes:

[0047] When the learner increases character input verification, the system comprehensively analyzes the speed and error rate during character input detection to obtain the accuracy score of the learner's character input. The speed during character input is measured by calculating the number of characters input per unit time. Common units are "words / minute" or "characters / second". The formula for the speed S of character input: S = ; Suppose, if someone inputs 300 characters in 1 minute, then the speed is 300 characters / minute; the error rate during character input is a common indicator used to evaluate text errors in the learner's character input. The error rate The calculation formula: = ; Then, the calculation formula for the correct rate W is: W = 1 - ; Where, S is the number of replaced incorrect characters, that is, the number of different characters between the characters input by the learner and the correct character output; D is the number of deleted incorrect characters, that is, the number of characters missed in the learner's character input; I is the number of inserted incorrect characters, that is, the number of extra characters in the learner's character input; N is the total number of characters in the reference text, that is, the number of characters in the correct text. The value of the error rate is usually between 0 and 1. The smaller the value, the lower the error rate of the learner's character input. Suppose the reference text is "hello" (N = 5), and the system output is "hxlxo". The calculation steps are as follows: Replaced errors (S): 'e' is replaced by 'x', 'l' is replaced by 'x', a total of 2 replaced errors; Deleted errors (D): No characters are deleted; Inserted errors (I): No extra characters are inserted; Therefore, the error rate : = = 0.4, which means the error rate is 40%, then, W = 1 - 0.4 = 0.6;

[0048] To comprehensively consider the speed and error rate during character input, we can represent the accuracy score of a learner's character input through a comprehensive scoring formula, combining the speed S and the error rate E into a weighted scoring model: Accuracy score : = ; where S is the speed, E is the error rate, and (1 - E) is the correct rate W, , are the weights of the speed and the correct rate respectively.

[0049] Merge the words in the text of the character input into the input set, and merge the words in the text after converting the speech data into the conversion set. Obtain the contrast of the character input by calculating the ratio of the intersection and union of the two sets. The contrast formula is: J = ; where, is the input set, and B is the conversion set; is the size of the intersection of the two sets, representing the number of words that appear together, is the size of the union of the two sets, representing the total number of all different words, and J is the contrast.

[0050] In this embodiment, S4 includes the following process:

[0051] By comprehensively analyzing the accuracy score and the contrast, evaluate the percentage of the accuracy of obtaining learning materials for character input. Use a weighted formula to calculate the percentage of the accuracy of obtaining G. The specific formula is: Percentage of accuracy G: G = ( ) 100%; where the accuracy score : = ; is the contrast J : J = ; is the weight of the accuracy score, is the weight of the contrast. Assume that the accuracy score = 0.85, and the contrast = 0.75. If the weight of the accuracy score = 0.6, and the weight of the contrast = 0.4, then the percentage of the accuracy of obtaining G is: G = ) 100% = (0.51 0.3) 100% = 81%, that is, when adding character input verification, the accuracy of obtaining learning materials is 81%.

[0052] Example 2: The present invention discloses a handwritten character recognition system based on a convolutional neural network, including a data acquisition module, an input selection module, a character analysis module, and a comprehensive evaluation module;

[0053] The data acquisition module is used to collect voice data, extract and quantify noise through sound pressure level, record the device temperature, extract the number of noise types, and analyze the change of the noise spectrum;

[0054] The input selection module is used to analyze the accuracy of voice input according to the sound pressure level and the change of the noise spectrum, and select the voice input method or add character input verification on the basis of the voice input method;

[0055] The character analysis module is used to detect the character input speed and error rate, analyze the accuracy score, and obtain the contrast by comparing the character input with the text converted from voice;

[0056] The comprehensive evaluation module is used to comprehensively analyze the accuracy score and the contrast, and evaluate the percentage of the accuracy of obtaining learning materials by character input.

[0057] Further, the data acquisition module includes a voice data acquisition sub-module and a noise processing sub-module;

[0058] The voice data acquisition sub-module is used to collect the voice data of the user and extract the noise therein; the noise processing sub-module is used to perform sound pressure level quantization processing on the noise extracted by the voice data acquisition sub-module, and record the temperature of the recording device and the number of noise types for subsequent spectrum change analysis;

[0059] The input selection module includes a noise analysis sub-module and a data input method selection sub-module;

[0060] The noise analysis sub-module analyzes the accuracy of obtaining learning materials through the voice input method based on the sound pressure level and spectrum change in the noise processing sub-module; the data input method selection sub-module selects different data input methods according to the accuracy obtained in the noise analysis sub-module, uses the voice input method for high accuracy, and adds character input verification on the basis of the voice input method for low accuracy;

[0061] The character analysis module includes a character input detection sub-module and a text comparison sub-module;

[0062] The character input detection sub-module is used to detect the speed and error rate of the user's character input, and analyze the accuracy score of the character input based on the speed and error rate;

[0063] The text comparison sub-module is used to convert the voice data into text, and compare the text of the character input with the text converted from voice to calculate the contrast of the character input;

[0064] The comprehensive evaluation module integrates the accuracy scores and contrast ratios in the character input detection sub-module and the text comparison sub-module, and comprehensively evaluates the percentage of the accuracy of character input for obtaining learning materials based on the integrated data.

[0065] In summary, the present invention can intelligently select the most suitable data input method and conduct a comprehensive evaluation based on the user's input performance, thereby improving the accuracy and efficiency of handwritten character recognition.

[0066] By obtaining the percentage of the accuracy of character input for obtaining learning materials, the accuracy of character input under different learning materials can be evaluated, effectively improving the efficiency of learners. Especially in scenarios where learning materials are diverse and high-precision input is required, it can effectively reduce the error rate of obtaining learning materials and improve learning outcomes.

[0067] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0068] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0069] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and invention constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0070] In addition, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0071] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0072] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A handwritten character recognition method based on a convolutional neural network, characterized in that; The steps include: S1. Collect voice data, extract noise from the voice data, and quantify the noise by sound pressure level; collect the temperature of the recording device when acquiring the voice data, and extract the number of noise types, and obtain the frequency spectrum change of the noise according to the temperature of the recording device and the number of noise types; S2. Analyze the accuracy of learning materials obtained through voice input through changes in sound pressure level and noise spectrum, select different data input methods according to the accuracy, use voice input for high accuracy, and add character input verification on top of voice input for low accuracy; S3, detecting the speed and error rate of character input, analyzing the accuracy score of character input; obtaining the contrast of character input by comparing the text of character input with the text converted from voice data; S4. Through comprehensive analysis of accuracy score and contrast, the accuracy percentage of character input in acquiring learning materials is evaluated.

2. The handwritten character recognition method based on convolutional neural network according to claim 1, characterized in that: The integrated temperature sensor of the recording device is used to collect the temperature of the recording device in real time when the voice data is acquired. By converting the voice data into two-dimensional spectrogram features, the convolutional neural network (CNN) can capture the noise features in the voice data and classify the noise based on this. The number of noise types is counted according to the noise features captured in the voice data. The speech data is processed in equal time segments to obtain multiple small audio segments, the sound pressure level of each small audio noise is quantified to obtain the noise sound pressure level, the spectrum of each small audio noise is processed to obtain the maximum spectrum point and the minimum spectrum point, and the spectrum change value of each small audio noise is obtained by subtracting the maximum spectrum point from the minimum spectrum point; The calculation formula for the sound pressure level of each small audio noise segment is: =20· ,in, is the sound pressure level; is the sound pressure at the measurement point; is the reference sound pressure; the calculation formula for the spectrum change value of each small audio noise is: ΔX = , where ΔX is the spectrum change value; the maximum spectrum point = , is the frequency corresponding to the maximum amplitude value, is the amplitude of the frequency point, the minimum spectrum point = , is the frequency corresponding to the minimum amplitude value, is the amplitude value at this frequency point, and are the maximum and minimum amplitude values ​​respectively.

3. The handwritten character recognition method based on convolutional neural network according to claim 2, characterized in that: The different noise sound pressure levels obtained from multiple short audio segments are merged into a sound pressure level dataset, and the different spectrum change values ​​obtained from multiple short audio segments are merged into a spectrum change dataset. When obtaining learning materials, the support vector machine model in the linear regression model is used to predict the accuracy of the learning materials obtained through voice input. The sound pressure level dataset and the spectrum change dataset are used as input variables to construct a support vector machine model. The support vector machine model is constructed using the sound pressure level dataset and the spectrum change dataset to predict the accuracy of the learning materials obtained through voice input.

4. The handwritten character recognition method based on convolutional neural network according to claim 3, characterized in that: The accuracy score of learners' character input is obtained by comprehensive analysis of the speed and error rate of character input. The speed of character input is measured by calculating the number of characters input per unit time. The speed of character input S is calculated as follows: S = ; Error rate The calculation formula is: = ; Then, the calculation formula of the accuracy rate W is: W = 1- ; Among them, S is the number of characters replaced incorrectly, that is, the number of different characters between the characters input by the learner and the correct character output; D is the number of characters deleted incorrectly, that is, the number of characters omitted by the learner's character input; I is the number of characters inserted incorrectly, that is, the number of extra characters in the learner's character input; N is the total number of characters in the reference text, that is, the number of characters in the correct text. The error rate value is usually between 0 and 1. The smaller the value, the lower the error rate of the learner's character input.

5. The handwritten character recognition method based on convolutional neural network according to claim 4, characterized in that: The comprehensive scoring formula is used to express the accuracy score of the learner's character input, combining the speed S and the error rate E into a weighted scoring model: Accuracy score : = ; Where S is the speed, E is the error rate, (1 - E) is the accuracy rate W, , are the weights of speed and accuracy respectively; The words in the text of character input are merged into the input set, and the words in the text converted from voice data are merged into the conversion set. The contrast of the character input is obtained by calculating the ratio of the intersection and union of the two sets. The contrast formula is: J = ;in, is the input set, B is the transformation set; is the size of the intersection of the two sets, indicating the number of words that appear together, is the union size of the two sets, indicating the total number of all different words, J It's the contrast.

6. The handwritten character recognition method based on convolutional neural network according to claim 5, characterized in that: The accuracy percentage G is calculated by a weighted formula. The specific formula is: Accuracy percentage G: G = ( ) 100%; among them, the accuracy score : = ; is the contrast J :J = ; is the weight of the precision score, is the contrast weight.

7. A handwritten character recognition system based on a convolutional neural network, used to implement the handwritten character recognition method based on a convolutional neural network according to any one of claims 1 to 6, characterized in that: It includes data acquisition module, input selection module, character analysis module and comprehensive evaluation module; Data acquisition module, used to collect voice data, extract and quantify noise through sound pressure level, record device temperature, extract the number of noise types, and analyze noise spectrum changes; An input selection module is used to analyze the accuracy of voice input according to the changes in sound pressure level and noise spectrum, and select a voice input mode or add character input verification based on the voice input mode; The character analysis module is used to detect the character input speed and error rate, analyze the accuracy score, and compare the character input with the speech-to-text to obtain contrast; A comprehensive evaluation module is used to detect character input speed and error rate, analyze accuracy scores, and compare character input with speech-to-text to obtain contrast.