Low-power intelligent screening method and system for learning disabilities based on multimodal data
By combining multiple modal data and specific processing models, the problems of low accuracy of learning disability screening and high consumption of computing resources in the prior art are solved, and a high-precision and low-consumption learning disability screening method is realized.
Patent Information
- Application Number
- CN202510156638.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art usually only uses single mode data in student learning disability screening, resulting in low accuracy of screening results and high consumption of computing resources, limiting its wide and in-depth application.
A low-consumable learning disability intelligent screening method based on multimodal data is adopted. By combining face images, voice signals and answering paper images, specific image matching, speech denoising and recognition models are used to achieve high accuracy and low consumption learning disability screening.
It significantly improves the accuracy of learning disability screening, reduces the consumption of computing resources, and enhances the practical application value of the method.
Smart Images

Figure CN119649435B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular, to a low-power intelligent screening method and system for learning disabilities based on multi-modal data. Background Art
[0002] With the development of the times and the progress of society, the education problems of students have attracted extensive attention from all sectors of society. However, there are still many students in the student group who have certain learning disabilities, which not only seriously affect their learning quality, but also easily cause more psychological problems.
[0003] Although some artificial intelligence technologies have been applied to the screening of students' learning disabilities, these technologies often only use single-modal data to obtain screening results, resulting in low accuracy of the screening results and reducing the practical application value of intelligent screening of learning disabilities. At the same time, the consumption of more computing resources further restricts the wide and in-depth application of intelligent screening of learning disabilities.
[0004] Therefore, it is of great value and significance to make full use of multi-modal data to achieve low-power screening of learning disabilities. Summary of the Invention
[0005] All actions of obtaining information or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization of the corresponding users.
[0006] In order to overcome the above problems or at least partially solve the above problems, the present invention provides a low-power intelligent screening method and system for learning disabilities based on multi-modal data, which makes full use of multi-modal data to achieve high-precision and low-power screening of learning disabilities.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] In a first aspect, the present invention provides a low-power intelligent screening method for learning disabilities based on multi-modal data, including the following steps:
[0009] Collect face images of students during the learning process according to a preset extraction time period;
[0010] Use a multi-level sparse coding image matching model based on Haar wavelet transform and a preset reference abnormal face image to detect whether each face image is abnormal, and generate a face detection result;
[0011] When the student's learning ends, obtain and use the student's question-and-answer voice signal as the voice signal to be detected;
[0012] Use a discriminative speech denoising model that combines Wiener filtering and Transformer to denoise the speech signal to be detected, obtaining the denoised speech signal to be detected;
[0013] Use a result evaluation-based speech recognition model that combines HMM and Transformer to recognize the denoised speech signal to be detected, and send the speech signal recognition result to the corresponding evaluator;
[0014] Obtain the evaluation result of the student's abnormal language expression by the evaluator;
[0015] When the student completes the voice Q&A, obtain the student's answer sheet image;
[0016] Use a clarity multi-dimensional analysis-based text recognition model that combines CNN and Transformer to recognize the student's answer sheet image, and send the answer sheet text recognition result to the corresponding evaluator;
[0017] Obtain the evaluation result of the student's abnormal writing expression by the evaluator;
[0018] Judge whether the student has a learning disorder based on the face detection result, the evaluation result of the student's abnormal language expression, and the evaluation result of the student's abnormal writing expression, and generate the final student learning disorder screening result.
[0019] First, the present invention proposes a multi-level sparse coding image matching model based on Haar wavelet transform to match face images. This model performs wavelet transform on the extracted face images and benchmark abnormal face images, conducts sparse coding and matching at multiple image levels, significantly improving the matching accuracy. Second, the present invention proposes a discriminative speech denoising model based on the combination of Wiener filter and Transformer to denoise speech signals. This model, based on the equal division detection of speech signals, selectively uses Wiener filter or Transformer to denoise each equal division of speech signals, not only ensuring the quality of speech denoising but also significantly reducing the consumption of computing resources. Third, the present invention proposes a result evaluation-based speech recognition model based on the combination of HMM and Transformer to recognize speech signals. This model first uses a simple HMM model to recognize speech signals and evaluates the recognition results. If the recognition results do not contain easily confused words, the recognized results are directly output; if the speech recognition results contain easily confused words, the Transformer model is used to recognize the speech signals again. This speech recognition method not only ensures the accuracy of speech recognition but also significantly reduces the consumption of computing resources. Finally, the present invention proposes a clarity multi-dimensional analysis-based text recognition model based on the combination of CNN and Transformer to recognize text. This model first analyzes the clarity of the image using various methods. If the evaluation results of each method indicate that the image clarity is high, the CNN model is directly used to recognize the text; otherwise, the Transformer model is used to recognize the text. This text recognition method not only ensures the accuracy of text recognition but also significantly reduces the consumption of computing resources.
[0020] In summary, the present invention makes full use of multi-modal data to achieve high-precision and low-consumption learning disorder screening, with good practical application value.
[0021] Based on the first aspect, further, the low-consumption learning disorder intelligent screening method based on multi-modal data further includes the following steps:
[0022] Use a reconstruction detection-based speech coding model based on multi-loss functions to encode the denoised speech signal to be detected, obtaining the encoded result of the high-quality denoised speech signal to be detected.
[0023] Based on the first aspect, further, the method of using a reconstruction detection-based speech coding model based on multi-loss functions to encode the denoised speech signal to be detected includes the following steps:
[0024] Use an autoencoder to encode the denoised speech signal to be detected, obtaining the encoded result; the loss function is used for constraint during the encoding process;
[0025] Reconstruct the encoded result into a speech signal, and detect the distortion degree of the reconstructed speech signal; if the distortion degree of the reconstructed speech signal is less than the preset distortion degree threshold, then use this encoded result as the encoded result of the speech signal to be detected after high-quality denoising; otherwise, continue to encode the speech signal to be detected after denoising using an autoencoder, and use a new loss function for constraint to obtain a new encoded result; reconstruct the new encoded result into a new speech signal, and detect the distortion degree of the reconstructed new speech signal; if the distortion degree of the new speech signal is less than the preset distortion degree threshold, then use this encoded result as the encoded result of the speech signal to be detected after high-quality denoising; otherwise, continue the above steps until the encoded result of the speech signal to be detected after high-quality denoising is obtained.
[0026] Based on the first aspect, further, the low-power learning disorder intelligent screening method based on multi-modal data further includes the following steps:
[0027] Obtain and upload the student identity information, face image, student answer sheet image, encoded result of the speech signal to be detected after high-quality denoising, and student learning disorder screening result to the blockchain for on-chain storage.
[0028] Based on the first aspect, further, the method for detecting whether there is an abnormality in each face image by using the multi-level sparse coding image matching model based on Haar wavelet transform and the preset reference abnormal face image includes the following steps:
[0029] Perform Haar wavelet transform on the face image and the reference abnormal face image;
[0030] For each level of the images after wavelet decomposition of the face image and the reference abnormal face image, perform sparse coding and matching respectively to obtain the matching degree of the corresponding level of images;
[0031] If the matching degree of each level of images is greater than the preset matching degree threshold, it is determined that the corresponding face image is abnormal; otherwise, it is determined that the corresponding face image is normal.
[0032] Based on the first aspect, further, the method for denoising the speech signal to be detected by using the discriminative speech denoising model combining Wiener filter and Transformer includes the following steps:
[0033] Perform equal division processing on the speech signal to be detected to obtain multiple segments of equally divided speech signals;
[0034] Perform peak signal-to-noise ratio detection on each segment of equally divided speech signal to obtain the peak signal-to-noise ratio of the corresponding equally divided speech signal;
[0035] If the peak signal-to-noise ratio of the equally divided speech signal is higher than the preset peak signal-to-noise ratio threshold, then use the Wiener filtering model to denoise the corresponding equally divided speech signal; otherwise, use the Transformer model to denoise the corresponding equally divided speech signal.
[0036] Based on the first aspect, further, the method for recognizing the denoised speech signal to be detected by using the result evaluation type speech recognition model combining HMM and Transformer includes the following steps:
[0037] Use the HMM model to recognize the denoised speech signal to be detected and generate a recognition result;
[0038] If the recognition result does not contain the preset easily confused words, then directly output the recognition result as the final speech signal recognition result; otherwise, use the Transformer model to recognize the denoised speech signal to be detected to obtain the final speech signal recognition result.
[0039] Based on the first aspect, further, the method for recognizing the student answer sheet image by using the clarity multi-dimensional analysis type character recognition model combining CNN and Transformer includes the following steps:
[0040] Use a variety of image clarity analysis methods to analyze the clarity of the student answer sheet image and generate corresponding multiple image clarity analysis results;
[0041] If all the image clarity analysis results are greater than the preset image clarity threshold, then use the CNN model to recognize the characters to obtain the answer sheet character recognition result; otherwise, use the Transformer model to recognize the characters to obtain the final answer sheet character recognition result.
[0042] Based on the first aspect, further, the above image clarity analysis methods include Brenner gradient method, Tenegrad gradient method, Laplace gradient method, variance method, and quantity gradient method.
[0043] In the second aspect, the present invention provides a low-power learning disorder intelligent screening system based on multi-modal data, including a face acquisition module, a face anomaly detection module, a voice acquisition module, a voice detection module, a voice recognition module, a language evaluation acquisition module, a drawing acquisition module, a character recognition module, a writing evaluation acquisition module, and a learning disorder judgment module, wherein:
[0044] The face acquisition module is used to acquire the face images of students during the learning process according to the preset extraction time period;
[0045] The face anomaly detection module is used to detect whether there is an anomaly in each face image by using a multi-level sparse coding image matching model based on Haar wavelet transform and a preset benchmark abnormal face image, and generate a face detection result;
[0046] The voice acquisition module is used to obtain and use the student's question-and-answer voice signal as the voice signal to be detected after the student finishes learning;
[0047] The voice detection module is used to denoise the voice signal to be detected by using a discriminative voice denoising model combining Wiener filtering and Transformer, and obtain the denoised voice signal to be detected;
[0048] The voice recognition module is used to recognize the denoised voice signal to be detected by using a result evaluation type voice recognition model combining HMM and Transformer, and send the voice signal recognition result to the corresponding evaluator;
[0049] The language evaluation acquisition module is used to obtain the evaluation result of the student's language expression anomaly by the evaluator;
[0050] The drawing acquisition module is used to obtain the student's answer sheet image after the student finishes the voice question and answer;
[0051] The character recognition module is used to recognize the student's answer sheet image by using a clarity multi-dimensional analysis type character recognition model combining CNN and Transformer, and send the answer sheet character recognition result to the corresponding evaluator;
[0052] The writing evaluation acquisition module is used to obtain the evaluation result of the student's writing expression anomaly by the evaluator;
[0053] The learning disorder judgment module is used to judge whether the student has a learning disorder according to the face detection result, the evaluation result of the student's language expression anomaly and the evaluation result of the student's writing expression anomaly, and generate the final screening result of the student's learning disorder.
[0054] The present invention has at least the following advantages or beneficial effects:
[0055] 1. The present invention proposes a multi-level sparse coding image matching model based on Haar wavelet transform to match face images; this model performs wavelet transform on the extracted face image and the benchmark abnormal face image, performs sparse coding and matching at multiple image levels, and significantly improves the matching accuracy.
[0056] 2. The present invention proposes a discriminative speech denoising model based on the combination of Wiener filtering and Transformer to denoise speech signals. Based on the equal division detection of speech signals, for each segment of equal division speech signal, Wiener filtering or Transformer is selectively used for denoising, which not only ensures the quality of speech denoising but also significantly reduces the consumption of computing resources.
[0057] 3. The present invention proposes a result evaluation type speech recognition model based on the combination of HMM and Transformer to recognize speech signals. The model first uses a simple HMM model to recognize speech signals and evaluates the recognition results. If the recognition results do not contain easily confused words, the recognized results are directly output. If the speech recognition results contain easily confused words, the Transformer model is used to recognize the speech signals again. This speech recognition method not only ensures the accuracy of speech recognition but also significantly reduces the consumption of computing resources.
[0058] 4. The present invention proposes a reconstruction detection type speech coding model based on multiple loss functions to code speech signals. Based on multiple different loss functions, the model selects high-quality coding results in the way of reconstruction detection, significantly improving the quality of speech coding.
[0059] 5. The present invention proposes a clarity multi-dimensional analysis type character recognition model based on the combination of CNN and Transformer to recognize characters. The model first analyzes the clarity of the image using various methods. If the evaluation results of each method are that the image clarity is high, the CNN model is directly used to recognize the characters. Otherwise, the Transformer model is used to recognize the characters. This character recognition method not only ensures the accuracy of character recognition but also significantly reduces the consumption of computing resources.
[0060] 6. The present invention utilizes blockchain technology to store the student's identity information, the extracted face image of the student, the answer sheet image, the coded result of the high-quality denoised speech signal to be detected, and the learning disorder screening result on the chain, ensuring the security of the system.
[0061] 7. The present invention makes full use of data of multiple modalities to achieve learning disorder screening with high accuracy and low consumption, having good practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0063] Figure 1 It is a flowchart of an intelligent screening method for low-consumption learning disabilities based on multi-modal data according to an embodiment of the present invention;
[0064] Figure 2 It is a principle block diagram of an intelligent screening system for low-consumption learning disabilities based on multi-modal data according to an embodiment of the present invention;
[0065] Figure 3 It is a structural block diagram of an electronic device provided by an embodiment of the present invention.
[0066] Explanation of reference numerals: 100, face acquisition module; 200, face anomaly detection module; 300, voice acquisition module; 400, voice detection module; 500, voice recognition module; 600, language evaluation acquisition module; 700, drawing acquisition module; 800, character recognition module; 900, writing evaluation acquisition module; 1000, learning disability judgment module; 101, memory; 102, processor; 103, communication interface. Detailed implementation manners
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations.
[0068] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0069] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0070] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0071] In the description of the embodiments of the present invention, "a plurality" represents at least two. Embodiment
[0072] As Figure 1 As shown, in a first aspect, an embodiment of the present invention provides a low-power learning disorder intelligent screening method based on multimodal data, including the following steps:
[0073] S1. According to a preset extraction time period, collect face images of students during the learning process; under the guidance of a tester, a certain student learns a specific learning content. During the learning process, the image acquisition device extracts the face images of the student at fixed time intervals.
[0074] S2. Use a multi-level sparse coding image matching model based on Haar wavelet transform and a preset reference abnormal face image to detect whether each face image has an abnormality, and generate a face detection result;
[0075] Further, it includes: performing Haar wavelet transform on the face image and the reference abnormal face image; performing sparse coding and matching on each level of the images after wavelet decomposition of the face image and the reference abnormal face image respectively to obtain the matching degree of the corresponding each level of image; if the matching degree of each level of image is greater than a preset matching degree threshold, it is determined that the corresponding face image has an abnormality; otherwise, it is determined that the corresponding face image is normal.
[0076] In some embodiments of the present invention, a multi-level sparse coding image matching model based on Haar wavelet transform is used to detect whether each face image has an abnormality. If most face images have an abnormality, it is determined that the facial expression of the student is abnormal. The above reference abnormal face image refers to the face image of a typical student with learning disorders selected.
[0077] S3. After the student finishes learning, obtain and use the student's question-and-answer voice signal as the voice signal to be detected;
[0078] S4. Use a discriminative speech denoising model that combines Wiener filtering and Transformer to denoise the speech signal to be detected, and obtain the denoised speech signal to be detected.
[0079] Further, it includes: equally dividing the speech signal to be detected to obtain multiple segments of equally divided speech signals; detecting the peak signal-to-noise ratio of each segment of equally divided speech signals to obtain the peak signal-to-noise ratio of the corresponding equally divided speech signals; if the peak signal-to-noise ratio of the equally divided speech signals is higher than the preset peak signal-to-noise ratio threshold, use the Wiener filtering model to denoise the corresponding equally divided speech signals; otherwise, use the Transformer model to denoise the corresponding equally divided speech signals.
[0080] Further, it also includes: using a reconstruction detection-based speech coding model based on multiple loss functions to encode the denoised speech signal to be detected, and obtaining the encoding result of the high-quality denoised speech signal to be detected. Specifically, it includes: using an autoencoder to encode the denoised speech signal to be detected to obtain an encoding result; using a loss function to constrain during the encoding process; reconstructing the encoding result into a speech signal, and detecting the distortion degree of the reconstructed speech signal; if the distortion degree of the reconstructed speech signal is less than the preset distortion degree threshold, use this encoding result as the encoding result of the high-quality denoised speech signal to be detected; otherwise, continue to use the autoencoder to encode the denoised speech signal to be detected, and use a new loss function to constrain to obtain a new encoding result; reconstruct the new encoding result into a new speech signal, and detect the distortion degree of the reconstructed new speech signal; if the distortion degree of the new speech signal is less than the preset distortion degree threshold, use this encoding result as the encoding result of the high-quality denoised speech signal to be detected; otherwise, continue the above steps until the encoding result of the high-quality denoised speech signal to be detected is obtained.
[0081] In some embodiments of the present invention, use an autoencoder to encode the denoised speech signal to be detected (constrained by a certain loss function during this process) to obtain an encoding result, reconstruct the encoding result into a speech signal. If there is no large distortion in the reconstructed speech signal, use this encoding result as the encoding result of the high-quality denoised speech signal to be detected. If there is a large distortion in the reconstructed speech signal, continue to use the autoencoder to encode the denoised speech signal to be detected (constrained by a new loss function during this process) to obtain an encoding result, reconstruct the encoding result into a speech signal. If there is no large distortion in the reconstructed speech signal, use this encoding result as the encoding result of the high-quality denoised speech signal to be detected. If there is a large distortion in the reconstructed speech signal, then continue the above process until the encoding result of the high-quality denoised speech signal to be detected is obtained.
[0082] S5. Use the result evaluation-based speech recognition model that combines HMM and Transformer to recognize the denoised speech signal to be detected, and send the speech signal recognition result to the corresponding evaluator;
[0083] Further, it includes: using the HMM model to recognize the denoised speech signal to be detected, and generating a recognition result; if the recognition result does not contain preset easily confused words, directly output the recognition result as the final speech signal recognition result; otherwise, use the Transformer model to recognize the denoised speech signal to be detected to obtain the final speech signal recognition result.
[0084] In some embodiments of the present invention, for the denoised speech signal to be detected, use the HMM model to recognize the speech signal. If the recognition result does not contain easily confused words, directly output the recognized result; if the recognition result contains easily confused words (such as the word 'poet', which has a highly similar pronunciation to words like 'ten people' and is prone to recognition errors), then use the Transformer model to recognize the speech signal again to obtain the final recognition result.
[0085] S6. Obtain the evaluation result of the student's abnormal language expression by the evaluator; a professional evaluator evaluates the recognition result of the denoised speech signal to be detected to determine whether the student has abnormal language expression, providing effective support for subsequent judgment of the student's learning disorder.
[0086] S7. After the student completes the voice Q&A, obtain the student's answer sheet image; after the student completes the voice answering questions session, under the guidance of the tester, the student answers relevant questions in writing on the answer sheet as required. After the student finishes writing the answer, use an image acquisition device to take a photo of the student's answer sheet to obtain the answer sheet image.
[0087] S8. Use the clarity multi-dimensional analysis-based text recognition model that combines CNN and Transformer to recognize the student's answer sheet image, and send the answer sheet text recognition result to the corresponding evaluator;
[0088] Further, it includes: analyzing the clarity of the student's answer sheet image using multiple image clarity analysis methods to generate corresponding multiple image clarity analysis results; if all the image clarity analysis results are greater than a preset image clarity threshold, using a CNN model to recognize the text to obtain the answer sheet text recognition result; otherwise, using a Transformer model to recognize the text to obtain the final answer sheet text recognition result. The above-mentioned image clarity analysis methods include Brenner gradient method, Tenegrad gradient method, Laplace gradient method, variance method, and quantity gradient method.
[0089] S9. Obtain the evaluation result of the student's abnormal writing expression by the evaluator; the professional personnel evaluate the answer sheet text recognition result to judge whether the student has abnormal writing expression.
[0090] S10. Judge whether the student has learning disabilities according to the face detection result, the evaluation result of the student's abnormal language expression, and the evaluation result of the student's abnormal writing expression, and generate the final student learning disability screening result.
[0091] In some embodiments of the present invention, if the student does not have abnormal facial expressions, abnormal language expressions, and abnormal writing expressions, it is determined that the student has no obvious learning disabilities; if the student has one of the abnormal facial expressions, abnormal language expressions, and abnormal writing expressions, it is determined that the student has general learning disabilities; if the student has more than two of the abnormal facial expressions, abnormal language expressions, and abnormal writing expressions, it is determined that the student has serious learning disabilities.
[0092] Further, the method further includes: obtaining and uploading the student identity information, face image, student answer sheet image, high-quality denoised speech signal coding result to be detected, and the student learning disability screening result to the blockchain to achieve on-chain storage. Note: The number of images is relatively small, so no coding is done; the voice is generally long, so coding is performed.
[0093] First, the present invention proposes a multi-level sparse coding image matching model based on Haar wavelet transform to match face images. This model performs wavelet transform on the extracted face images and reference abnormal face images, conducts sparse coding and matching at multiple image levels, significantly improving the matching accuracy. Secondly, the present invention proposes a discriminative speech denoising model based on the combination of Wiener filter and Transformer to denoise speech signals. This model, based on the equal division detection of speech signals, selectively uses Wiener filter or Transformer to denoise each equal division of speech signals, not only ensuring the quality of speech denoising but also significantly reducing the consumption of computing resources. Thirdly, the present invention proposes a result evaluation-based speech recognition model based on the combination of HMM and Transformer to recognize speech signals. This model first uses a simple HMM model to recognize speech signals and evaluates the recognition results. If the recognition results do not contain easily confused words, the recognition results are directly output. If the speech recognition results contain easily confused words, the Transformer model is used to recognize the speech signals again. This speech recognition method not only ensures the accuracy of speech recognition but also significantly reduces the consumption of computing resources. Then, the present invention proposes a reconstruction detection-based speech coding model based on multiple loss functions to code speech signals. This model, based on multiple different loss functions, selects high-quality coding results in a reconstruction detection manner, significantly improving the quality of speech coding. Subsequently, the present invention proposes a clarity multi-dimensional analysis-based text recognition model based on the combination of CNN and Transformer to recognize text. This model first analyzes the clarity of the image using various methods. If the evaluation results of each method indicate that the image clarity is high, the CNN model is directly used to recognize the text. Otherwise, the Transformer model is used to recognize the text. This text recognition method not only ensures the accuracy of text recognition but also significantly reduces the consumption of computing resources. Finally, the present invention utilizes blockchain technology to store the student's identity information, the extracted face image of this student, the answer sheet image, the coding results of the high-quality denoised speech signals to be detected, and the learning disorder screening results on the chain, ensuring the security of the system. In summary, the present invention makes full use of various modal data to achieve high-precision and low-consumption learning disorder screening, having good practical application value.
[0094] Such as Figure 2As shown in the figure, in a second aspect, an embodiment of the present invention provides a low-power intelligent screening system for learning disabilities based on multi-modal data, including a face acquisition module 100, a face anomaly detection module 200, a voice acquisition module 300, a voice detection module 400, a voice recognition module 500, a language evaluation acquisition module 600, a drawing acquisition module 700, a text recognition module 800, a writing evaluation acquisition module 900, and a learning disability judgment module 1000, where:
[0095] The face acquisition module 100 is used to collect face images of students during the learning process according to a preset extraction time period;
[0096] The face anomaly detection module 200 is used to detect whether there is an anomaly in each face image by using a multi-level sparse coding image matching model based on Haar wavelet transform and a preset reference abnormal face image, and generate a face detection result;
[0097] The voice acquisition module 300 is used to obtain and use the student's question-and-answer voice signal as the voice signal to be detected after the student's learning ends;
[0098] The voice detection module 400 is used to denoise the voice signal to be detected by using a discriminative voice denoising model based on the combination of Wiener filtering and Transformer to obtain the denoised voice signal to be detected;
[0099] The voice recognition module 500 is used to recognize the denoised voice signal to be detected by using a result evaluation type voice recognition model based on the combination of HMM and Transformer, and send the voice signal recognition result to the corresponding evaluator;
[0100] The language evaluation acquisition module 600 is used to obtain the evaluation result of the student's language expression anomaly by the evaluator;
[0101] The drawing acquisition module 700 is used to obtain the student's answer sheet image after the student completes the voice question and answer;
[0102] The text recognition module 800 is used to recognize the student's answer sheet image by using a clarity multi-dimensional analysis type text recognition model based on the combination of CNN and Transformer, and send the answer sheet text recognition result to the corresponding evaluator;
[0103] The writing evaluation acquisition module 900 is used to obtain the evaluation result of the student's writing expression anomaly by the evaluator;
[0104] The learning disability judgment module 1000 is used to judge whether the student has a learning disability according to the face detection result, the evaluation result of the student's language expression anomaly, and the evaluation result of the student's writing expression anomaly, and generate the final screening result of the student's learning disability.
[0105] Through the cooperation of multiple modules such as the face acquisition module 100, the face anomaly detection module 200, the voice acquisition module 300, the voice detection module 400, the voice recognition module 500, the language evaluation acquisition module 600, the drawing acquisition module 700, the character recognition module 800, the writing evaluation acquisition module 900, and the learning disability judgment module 1000, this system makes full use of data of multiple modalities to achieve learning disability screening with relatively high accuracy and low consumption, and has good practical application value. First, the present invention proposes a multi-level sparse coding image matching model based on Haar wavelet transform to match face images; this model performs wavelet transform on the extracted face image and the reference abnormal face image, performs sparse coding and matching at multiple image levels, and significantly improves the accuracy of matching. Second, the present invention proposes a discriminative speech denoising model based on the combination of Wiener filter and Transformer to denoise speech signals; this model, based on the equal division detection of speech signals, selectively uses Wiener filter or Transformer to denoise each equal division of speech signals, not only ensuring the quality of speech denoising, but also significantly reducing the consumption of computing resources. Third, the present invention proposes a result evaluation type speech recognition model based on the combination of HMM and Transformer to recognize speech signals; this model first uses a simple HMM model to recognize speech signals and evaluates the recognition results. If the recognition results do not contain easily confused words, the recognized results are directly output; if the speech recognition results contain easily confused words, the Transformer model is used to recognize the speech signals again. This speech recognition method not only ensures the accuracy of speech recognition, but also significantly reduces the consumption of computing resources. Finally, the present invention proposes a clarity multi-dimensional analysis type character recognition model based on the combination of CNN and Transformer to recognize characters; this model first analyzes the clarity of the image using multiple methods. If the evaluation results of each method are that the image clarity is relatively high, the CNN model is directly used to recognize the characters; otherwise, the Transformer model is used to recognize the characters. This character recognition method can not only ensure the accuracy of character recognition, but also significantly reduce the consumption of computing resources.
[0106] As Figure 3 shown, in the third aspect, an embodiment of the present application provides an electronic device, which includes a memory 101 for storing one or more programs; a processor 102. When the one or more programs are executed by the processor 102, the method as described in any item of the first aspect above is implemented.
[0107] It further includes a communication interface 103, and the memory 101, the processor 102, and the communication interface 103 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used to communicate signals or data with other node devices.
[0108] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0109] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0110] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the methods, systems, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0111] In addition, in each embodiment of this application, the various functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0112] Fourthly, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 102, it implements the method according to any one of the above first aspects. If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0113] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0114] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A low-cost intelligent screening method for learning disabilities based on multimodal data, characterized in that: The following steps are involved: According to the preset extraction time period, collect the face images of students during the learning process; Using a multi-level sparse coding image matching model based on Haar wavelet transform and a preset reference abnormal face image, each face image is detected to see if it is abnormal, and a face detection result is generated, including: performing Haar wavelet transform on the face image and the reference abnormal face image; performing sparse coding and matching on each level of the face image and the reference abnormal face image after wavelet decomposition, so as to obtain the matching degree of each level of the corresponding image; if the matching degree of each level of the image is greater than the preset matching degree threshold, the corresponding face image is considered to be abnormal; otherwise, the corresponding face image is considered to be normal; When the student finishes learning, the student's question and answer voice signal is obtained and used as the voice signal to be detected; The discriminative speech denoising model based on the combination of Wiener filtering and Transformer is used to denoise the speech signal to be detected, and the denoised speech signal to be detected is obtained; Using the result-evaluation speech recognition model based on the combination of HMM and Transformer, the denoised speech signal to be detected is recognized, and the speech signal recognition result is obtained and sent to the corresponding evaluator; Obtain the assessment results of students’ language expression abnormality from the assessor; When the student completes the voice question and answer, obtain the student answer sheet image; Using a clarity multi-dimensional analysis text recognition model based on a combination of CNN and Transformer, the student answer sheet image is recognized, and the answer sheet text recognition result is obtained and sent to the corresponding evaluator, including: using multiple image clarity analysis methods to analyze the clarity of the student answer sheet image, and generating multiple corresponding image clarity analysis results; if all the image clarity analysis results are greater than the preset image clarity threshold, the CNN model is used to recognize the text to obtain the answer sheet text recognition result; otherwise, the Transformer model is used to recognize the text to obtain the final answer sheet text recognition result; Obtain the assessment results of students’ written expression abnormalities from the assessors; Based on the face detection results, the student's language expression abnormality assessment results and the student's writing expression abnormality assessment results, it is judged whether the student has learning disabilities and the final student learning disability screening results are generated.
2. According to claim 1, a low-cost intelligent screening method for learning disabilities based on multimodal data is characterized in that: The following steps are also included: The denoised speech signal to be detected is encoded using a reconstruction detection speech coding model based on multiple loss functions to obtain a high-quality denoised speech signal encoding result.
3. According to claim 2, a low-cost intelligent screening method for learning disabilities based on multimodal data is characterized in that: The method for encoding the denoised speech signal to be detected by using a reconstruction detection speech coding model based on multiple loss functions comprises the following steps: The de-noised speech signal to be detected is encoded by using an automatic encoder to obtain an encoding result; the loss function is used to constrain the encoding process; The encoding result is reconstructed into a speech signal, and the distortion of the reconstructed speech signal is detected; if the distortion of the reconstructed speech signal is less than a preset distortion threshold, the encoding result is used as the encoding result of the high-quality denoised speech signal to be detected; otherwise, the denoised speech signal to be detected is continuously encoded using the automatic encoder, and a new loss function is used for constraint to obtain a new encoding result; the new encoding result is reconstructed into a new speech signal, and the distortion of the reconstructed new speech signal is detected; if the distortion of the new speech signal is less than a preset distortion threshold, the encoding result is used as the encoding result of the high-quality denoised speech signal to be detected; otherwise, the above steps are continued until the encoding result of the high-quality denoised speech signal to be detected is obtained.
4. According to claim 2, a low-cost intelligent screening method for learning disabilities based on multimodal data is characterized in that: The following steps are also included: Obtain and upload student identity information, facial images, student answer sheet images, high-quality denoised speech signal encoding results to be tested, and student learning disability screening results to the blockchain for on-chain storage.
5. According to claim 1, a low-cost intelligent screening method for learning disabilities based on multimodal data is characterized in that: The method for denoising the speech signal to be detected by using a distinguishable speech denoising model based on a combination of Wiener filtering and Transformer comprises the following steps: The speech signal to be detected is processed in equal parts to obtain a plurality of equally divided speech signals; Performing peak signal-to-noise ratio detection on each equally divided speech signal to obtain the peak signal-to-noise ratio of the corresponding equally divided speech signal; If the peak signal-to-noise ratio of the equally divided speech signal is higher than the preset peak signal-to-noise ratio threshold, the Wiener filter model is used to perform speech denoising on the corresponding equally divided speech signal; otherwise, the Transformer model is used to perform speech denoising on the corresponding equally divided speech signal.
6. The low-cost intelligent screening method for learning disabilities based on multimodal data according to claim 1, characterized in that: The method for recognizing the denoised speech signal to be detected by using the result evaluation speech recognition model based on the combination of HMM and Transformer comprises the following steps: Use the HMM model to recognize the denoised speech signal to be detected and generate a recognition result; If the recognition result does not contain the preset easily confused words, the recognition result is directly output as the final speech signal recognition result; otherwise, the Transformer model is used to recognize the denoised speech signal to be detected to obtain the final speech signal recognition result.
7. The low-cost intelligent screening method for learning disabilities based on multimodal data according to claim 1, characterized in that: The image clarity analysis methods include Brenner gradient method, Tenegrad gradient method, Laplace gradient method, variance method and quantity gradient method.
8. A low-cost intelligent screening system for learning disabilities based on multimodal data, characterized in that: It includes a face acquisition module, a face anomaly detection module, a voice acquisition module, a voice detection module, a voice recognition module, a language assessment acquisition module, a drawing acquisition module, a text recognition module, a writing assessment acquisition module and a learning disability judgment module, among which: The face collection module is used to collect face images of students during their learning process according to a preset extraction time period; The face anomaly detection module is used to detect whether each face image has an abnormality by using a multi-level sparse coding image matching model based on Haar wavelet transform and a preset reference abnormal face image, and generate a face detection result, including: performing Haar wavelet transform on the face image and the reference abnormal face image; performing sparse coding and matching on each level of the image after wavelet decomposition of the face image and the reference abnormal face image, respectively, to obtain the matching degree of each level of the corresponding image; if the matching degree of each level of the image is greater than the preset matching degree threshold, it is determined that the corresponding face image has an abnormality; otherwise, it is determined that the corresponding face image is normal; The voice acquisition module is used to obtain the student's question and answer voice signal as the voice signal to be detected after the student finishes learning; The speech detection module is used to denoise the speech signal to be detected by using a discriminative speech denoising model based on the combination of Wiener filtering and Transformer to obtain the denoised speech signal to be detected; The speech recognition module is used to recognize the denoised speech signal to be detected by using a result-evaluated speech recognition model based on a combination of HMM and Transformer, and obtain and send the speech signal recognition result to the corresponding evaluator; The language assessment acquisition module is used to obtain the assessment results of abnormal language expression of students from the assessor; The drawing acquisition module is used to obtain the student answer sheet image after the student completes the voice question and answer; The text recognition module is used to recognize the student answer sheet image using a clarity multi-dimensional analysis text recognition model based on a combination of CNN and Transformer, and obtain and send the answer sheet text recognition results to the corresponding evaluator, including: using multiple image clarity analysis methods to analyze the clarity of the student answer sheet image to generate multiple corresponding image clarity analysis results; if all the image clarity analysis results are greater than the preset image clarity threshold, then using the CNN model to recognize the text and obtain the answer sheet text recognition results; otherwise, using the Transformer model to recognize the text and obtain the final answer sheet text recognition results; A writing assessment acquisition module is used to obtain the assessment results of abnormal writing expressions of students from the assessors; The learning disability judgment module is used to judge whether a student has a learning disability based on the face detection results, the student's language expression abnormality assessment results and the student's writing expression abnormality assessment results, and generate the final student learning disability screening results.
Citation Information
Patent Citations
Learning input degree resource optimization type accurate detection method and system based on audio and video
CN117894303A