AR-AI integrated interactive learning assistance method

By integrating AR-AI into an interactive learning aid method, and utilizing image recognition and voice command technologies, the problem of insufficient accuracy in recognizing overlapping characters is solved, achieving a highly efficient learning aid effect.

CN120808648AActive Publication Date: 2025-10-17延边教育出版社
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511000214.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-11
Filing Date
2025-07-21
Publication Date
2025-10-17
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing image recognition technology lacks accuracy when text and images overlap, hindering its application and development in complex image environments.

Method used

An AR-AI integrated interactive learning assistance method is adopted, which uses technologies such as image recognition, feature extraction, voice command recognition and clustering algorithms to accurately identify overlapping text and provides learning assistance in combination with knowledge graphs.

Benefits of technology

It improves the accuracy and efficiency of recognition in cases of overlapping text, enhances the user's learning experience, and enables effective guidance for users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808648A_ABST
    Figure CN120808648A_ABST
Patent Text Reader

Abstract

The invention provides an AR-AI integrated interactive learning assistance method, and relates to the technical field of intelligent teaching. The method comprises the following steps: acquiring an image containing question content and segmenting the image to obtain at least one segmented region; acquiring character line data in the at least one segmented region through an image recognition technology; independently extracting character lines with different characteristics; acquiring a voice control instruction and selecting at least one text content for recognition and analysis; on the basis of a basic text recognition algorithm, the text recognition result is subjected to spot check, the correctness of the text recognition result is judged, and according to the correctness of the text recognition result, an improved image recognition algorithm is adopted for re-recognition, so that the correctness of image recognition can be ensured, and the learning efficiency of a user is improved when the user is assisted in learning. The text needing to be input by the user is accurately recognized, the recognition accuracy and efficiency are guaranteed, the use experience of the user is improved, and effective tutoring is conducted on the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent teaching technology, and in particular to an AR-AI integrated interactive learning assistance method. Background Art

[0002] With the rapid development of mobile Internet technology, photo recognition technology has been widely used in various fields as an efficient and convenient way to obtain information. In many scenarios such as education, life services, and medical consultation, applications based on photo recognition capture the image content taken by users and use advanced image processing and natural language processing technologies to quickly identify and analyze the information in the image, and then provide users with corresponding answers or solutions. These applications have greatly enriched the channels for information acquisition and improved the efficiency of problem solving.

[0003] In complex image environments, especially when text and images overlap, the accuracy of existing photo recognition technology is often seriously affected. Since the text in the image may be interfered with by various factors such as background patterns, colors, and lighting conditions, it is difficult for the recognition algorithm to accurately distinguish and extract valid text information. In addition, the close overlap of text and images may also cause confusion in the recognition area, further increasing the difficulty and error rate of recognition. This lack of recognition accuracy in the case of text and image overlap directly limits the application and development of photo recognition technology in a wider range of scenarios.

[0004] In view of the problems existing in existing photo recognition applications in terms of interactivity and recognition accuracy, especially the recognition difficulties in the case of overlapping text and images, it is necessary to carry out technical innovation and optimization. Therefore, the present invention provides an improved learning assistance system. Through the improved text recognition algorithm and process, even if the user records notes on text carriers such as books and there is text overlap with other text content, accurate recognition of text content can still be achieved, ensuring recognition accuracy and efficiency, improving the user experience, and providing effective guidance to users. Summary of the Invention

[0005] In order to overcome the problem of low recognition success rate and accuracy of overlapping text in the existing text recognition process, the present invention provides an AR-AI integrated interactive learning assistance method.

[0006] In order to solve the above technical problems, the technical solutions provided by the present invention are as follows:

[0007] The AR-AI integrated interactive learning assistance method includes the following steps:

[0008] S11: Acquire an image containing the topic content and segment the image to obtain at least one segmented region;

[0009] S12: Recognize the text lines in the at least one segmented region through image recognition technology to obtain text line data in the at least one segmented region;

[0010] S13: Extract features of text lines with different features separately through feature recognition technology and classification technology to obtain at least one text content;

[0011] S14: Obtain voice instruction information data of a user and recognize the voice instruction data of the user to obtain a voice control instruction;

[0012] S15: Select at least one text content for recognition and analysis according to a recognition result of the voice instruction data;

[0013] S16: Provide auxiliary guidance for learning of the user in at least one of a voice form, a text form, a video form and an image form according to a recognition and analysis result of the text content.

[0014] In any of the above solutions, preferably, when the text lines in the at least one segmented region are recognized through the image recognition technology to obtain the text line data in the segmented region, the following steps are included:

[0015] S21: Image preprocessing, enhancing text information in the image to make the text clearer and more identifiable;

[0016] S22: Feature extraction, extracting features of the text in the image through edge detection technology and texture analysis technology to obtain comprehensive text feature data;

[0017] S23: Image segmentation, finding an outline of the text by analyzing a connected domain in the image and segmenting according to text line features;

[0018] S24: Image post-processing, further optimizing and adjusting a segmentation result to improve accuracy and stability of recognition.

[0019] In any of the above solutions, preferably, when the text lines with different features are extracted separately through the feature recognition technology and the classification technology to obtain at least one text content, the following steps are included:

[0020] S31: Text recognition, recognizing a text based on the obtained comprehensive text feature data to obtain a first preset text;

[0021] S32: Randomly sampling the first preset text to sample 20%-40% of characters in the first preset text;

[0022] S33: When the sampling result is qualified, ending the text recognition, and when the sampling result is unqualified, proceeding to the next step;

[0023] S34: first, the text lines in the at least one segmentation region are separated by a color-based separation technique, if successful separation, at least two color-separated text image data are obtained, and text recognition is performed on the at least two separated text image data, then step S32 is performed for inspection, and if the inspection is qualified, the text recognition is ended, otherwise the next step is entered;

[0024] S35: image features are extracted and classified by using a clustering algorithm, and text images are separated by using morphological operations, at least two feature-separated text image data are obtained, and text recognition is performed on the at least two feature-separated text image data, then step S32 is performed for inspection, and if the inspection is qualified, the text recognition is ended, otherwise the clustering parameters of the clustering algorithm are modified, and the recognition is performed again.

[0025] In any of the above schemes, it is preferred that when 20%-40% of the characters in the first preset text are randomly sampled for inspection, the following steps are included:

[0026] S41: reference text selection, text line data in the segmentation region are recognized by an optical character recognition technique, a text font corresponding to the text line data is recognized, and a corresponding reference text is selected according to the text content in the first preset text;

[0027] S42: threshold value selection, different similarity thresholds are selected as the first threshold value for judging the correctness of the characters according to the types of the recognized text fonts, and a second threshold value for judging the text sampling result is calculated according to the image clarity, the font type, and the image background complexity;

[0028] S42: similarity calculation, characters in the first preset text are randomly selected, the number of the selected characters is 20%-40% of all the characters in the first preset text, and similarity calculation is performed on the images corresponding to the selected characters and the images of the corresponding reference text;

[0029] S43: result judgment, whether the characters are correct is judged according to the calculated similarity value and the selected first threshold value;

[0030] S44: comprehensive judgment, the judgment results of all the selected characters are comprehensively judged, the total accuracy is calculated, and whether the sampling result is qualified is judged according to the obtained second threshold value.

[0031] In any of the above schemes, it is preferred that when image features are extracted and classified by using a clustering algorithm, and text images are separated by using morphological operations, at least two feature-separated text image data are obtained, the following steps are included:

[0032] S51: high-level feature extraction, performing high-level feature extraction on the at least one segmented region, wherein the high-level feature extraction includes local binary pattern and gray level co-occurrence matrix;

[0033] S52: setting the number of clusters K = n + 2, wherein n is the number of times of detecting the segmented region, and inputting the extracted high-level feature and the comprehensive character feature data into a K-means clustering algorithm, and using the K-means clustering algorithm to divide the pixels and the pixel blocks into n + 2 clusters;

[0034] S53: removing the cluster represented by the background image, mapping the remaining clusters to the segmented region, and separating the character region and removing background noise through morphological operation to obtain at least two feature-separated character image data.

[0035] In any of the above schemes, preferably, when the character region is separated and the background noise is removed through morphological operation to obtain at least two feature-separated character image data, the following method is included:

[0036] A11: reducing the highlight area in the image through erosion operation, removing isolated noise points, and separating touching and approaching objects;

[0037] A12: expanding the highlight area in the image through dilation operation, filling the holes inside the object, and smoothing the object edge;

[0038] A13: removing small objects through opening operation, separating objects at thin points, and smoothing the boundary of larger objects;

[0039] A14: filling small holes inside the foreground object through closing operation, connecting adjacent objects, and smoothing the boundary thereof;

[0040] A15: highlighting the boundary of the object through morphological gradient operation.

[0041] In any of the above schemes, preferably, after text recognition on the at least two feature-separated character image data, the following steps are further included:

[0042] S61: position-based judgment, analyzing the position of the text in the image and matching it with the expected category or region;

[0043] S62: checking whether the text content is consistent with the image context;

[0044] S63: analyzing the semantic meaning of the text, and judging whether it is consistent with the overall theme and context of the image text;

[0045] S64: The original image of the identified error character is extracted separately, and the possible meaning of the error character is judged according to the semantic meaning of the text, and the original image of the error character is re-recognized according to the possible meaning of the error character as the judgment standard.

[0046] In any of the above schemes, preferably, in the process of obtaining the voice instruction information data of the user and recognizing the voice instruction data of the user, obtaining the voice control instruction comprises the following steps:

[0047] S71: Voice signal conversion and preprocessing, converting the collected voice signal into a digital signal, and filtering, pre-emphasizing, framing and windowing the obtained digital signal;

[0048] S72: Voice feature extraction, converting the pre-processed digital signal from time domain to frequency domain, and extracting parameters representing the essential features of the voice signal;

[0049] S73: Acoustic modeling, using a hidden Markov model to match the extracted feature parameters with a pre-set acoustic model;

[0050] S74: Language modeling and decoding, modeling language characteristics through a recurrent neural network language model, and then using a decoder to fuse the information of the acoustic model and the language model to parse the most likely text sequence.

[0051] In any of the above schemes, preferably, in the process of selecting at least one text content for recognition and analysis according to the recognition result of the voice instruction data, the following steps are included:

[0052] S81: Text preprocessing, preprocessing the parsed text sequence, wherein the preprocessing method includes noise removal, word segmentation and part-of-speech tagging;

[0053] S82: Named entity recognition, performing named entity recognition based on text preprocessing to identify specific entities in the instruction;

[0054] S83: Intent recognition, analyzing the semantic content of the instruction through natural language processing technology to identify the user's intent;

[0055] S84: Parameter extraction, further extracting key parameters in the instruction based on the identified user intent;

[0056] S85: Instruction parsing and construction, constructing an instruction parsing tree or similar data structure according to the user's intent and extracted parameters, and decomposing complex instructions into a series of executable sub-tasks;

[0057] S86: Execute the instruction.

[0058] In any of the above schemes, preferably, when providing auxiliary guidance for the user's learning through at least one of voice form, text form, video form and image form according to the recognition and analysis results of the text content, the following steps are included:

[0059] S91: According to the user's instruction, corresponding operation and answer are carried out;

[0060] S92: The knowledge points contained in the question content are deeply analyzed by using knowledge graph technology, and the key knowledge points are recognized and extracted, and the user's learning history and performance are predicted by using machine learning algorithm, and the difficulties that the user may encounter are predicted;

[0061] S93: According to the obtained knowledge points and difficulties, search in the pre-set database, search related audio, picture and video, and according to the set importance coefficient and the proportion coefficient of knowledge points in the question, sort the searched audio, picture and video, and take the order as the default playing order, and play to the user.

[0062] Compared with the prior art, the beneficial effects of the present application are:

[0063] The above scheme of the present application can ensure the correctness of image recognition by judging the correctness of the text recognition result on the basis of the basic text recognition algorithm, and re-recognizing by using the improved image recognition algorithm according to the correctness of the text recognition result, so as to accurately recognize the text needed to be input by the user when assisting the user to learn, ensure the accuracy and efficiency of recognition, increase the use experience of the user, and effectively guide the user.

[0064] When recognizing the overlapped text, the image is processed by using the clustering algorithm, the clustering algorithm is used for pretreatment and auxiliary separation, the text recognition algorithm is used for text recognition, the qualifiedness of the recognition result is detected, and the clustering number of the clustering algorithm is changed according to the detection result, and then the recognition is carried out again after the change, so as to effectively ensure the accuracy of the recognition result, and solve the problems of low recognition success rate and accuracy of overlapped text in the existing text recognition process. BRIEF DESCRIPTION OF DRAWINGS

[0065] Fig. 1 is a flowchart of an AR-AI integrated interactive learning auxiliary method provided by an embodiment of the present application.

[0066] Fig. 2 is a flowchart of text line extraction in an AR-AI integrated interactive learning auxiliary method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0068] Please refer to Figs. 1-2 The present application provides an example: AR-AI integrated interactive learning assistance method, including the following steps:

[0069] S11: obtaining an image containing the content of the question and segmenting the image to obtain at least one segmented region, which can specifically include:

[0070] Collecting an image of a paper containing the content of the question and segmenting the paper image to obtain at least one segmented region with independent text content;

[0071] S12: identifying the text lines in at least one segmented region by image recognition technology to obtain text line data in at least one segmented region, which can specifically include:

[0072] First, image preprocessing is performed to enhance the text information in the image and make the text clearer and more recognizable; then, feature extraction is performed on the preprocessed image, the features of the text in the image are extracted through edge detection technology and texture analysis technology, and comprehensive text feature data is obtained; image segmentation is performed according to the feature extraction result, the contour of the text is found by analyzing the connected domain in the image, and segmentation is performed according to the text line features; after segmentation, image post-processing is performed to further optimize and adjust the segmentation result, so as to improve the accuracy and stability of the recognition;

[0073] S13: extracting features of different text lines separately by feature recognition technology and classification technology to obtain at least one text content, which can specifically include:

[0074] First, text recognition is performed, and based on the obtained comprehensive text feature data, the text is recognized to obtain a first preset text; after text recognition, detection is performed, and 20%-40% of the characters in the first preset text are randomly sampled for inspection; when the sampling result is qualified, the text recognition is ended, and when the sampling result is unqualified, deep analysis is performed: first, the text lines in at least one segmentation region are separated based on a color-based separation technology, if the separation is successful, at least two color-separated text image data are obtained, and text recognition is performed on the at least two separated text image data, and inspection is performed again, and the text recognition is ended after the inspection is qualified, if at least two color-separated text image data cannot be separated or the inspection is unqualified, the image is recognized again through a clustering algorithm: the image features are extracted and classified by using the clustering algorithm, and the text image is separated by using a morphological operation to obtain at least two feature-separated text image data, and text recognition is performed on the at least two feature-separated text image data, and inspection is performed again, and the text recognition is ended after the inspection is qualified, otherwise, the clustering parameters of the clustering algorithm are modified, and the recognition is performed again;

[0075] S14: Obtain voice instruction information data of the user, and recognize the voice instruction data of the user to obtain a voice control instruction;

[0076] S15: According to the recognition result of the voice instruction data, at least one text content is selected for recognition and analysis;

[0077] S16: According to the recognition and analysis result of the text content, at least one of a voice form, a text form, a video form and an image form is used to provide auxiliary guidance for the learning of the user.

[0078] In the embodiment, by sampling the text recognition result based on the basic text recognition algorithm, the correctness of the text recognition result is judged, and according to the correctness of the text recognition result, the improved image recognition algorithm is used for re-recognition, which can ensure the correctness of the image recognition, so as to accurately recognize the text input by the user when assisting the user in learning, ensure the accuracy and efficiency of the recognition, increase the use experience of the user, and effectively guide the user.

[0079] It should be pointed out that the above steps are only preferred implementation sequences, and in the specific implementation process, part of the steps can be exchanged without affecting the overall implementation effect, and the following content is explained in a preferred manner to explain the technical scheme of the present application.

[0080] In the embodiment of the present application, when 20%-40% of the characters in the first preset text are randomly sampled for inspection, the following steps are included:

[0081] S41: reference text selection, recognizing the character trace data in the segmented region by optical character recognition technology, identifying the text font corresponding to the character trace data, and selecting the corresponding reference text according to the character content in the first preset text;

[0082] S42: threshold value selection, selecting different similarity thresholds as the first threshold value for judging the correctness of the characters according to the types of the identified text font, and calculating the second threshold value for judging the text sampling result according to the image clarity, font type and image background complexity;

[0083] S42: similarity calculation, randomly selecting characters in the first preset text, the number of selected characters being 20%-40% of all characters in the first preset text, and calculating the similarity between the image corresponding to the selected characters and the image of the corresponding reference text;

[0084] S43: result judgment, judging whether the characters are correct according to the calculated similarity value and the selected first threshold value;

[0085] S44: comprehensive judgment, synthesizing the judgment results of all selected characters, calculating the total accuracy, and judging whether the sampling result is qualified according to the obtained second threshold value.

[0086] Through the above steps S41-S44, in this embodiment, by combining optical character recognition (OCR) technology, the text font is automatically recognized and the corresponding reference text is selected, then 20%-40% of the characters randomly selected are calculated for image similarity, and the double threshold values are dynamically set according to the font type, image clarity and background complexity for correctness judgment. This method not only realizes the automatic monitoring of text quality, but also significantly improves the accuracy and efficiency of sampling, provides reliable quality guarantee for text processing process, and helps to continuously optimize the OCR algorithm and sampling strategy.

[0087] In another optional embodiment of the present application, when different similarity thresholds are selected as the first threshold value for judging the correctness of the characters according to the types of the identified text font, the similarity thresholds selected for different fonts are different, such as the similarity threshold for handwritten text being selected as 72%, the similarity threshold for Simplified Chinese Songti text being selected as 89%, and the similarity threshold for Traditional Chinese Songti text being selected as 81%. Through a large number of experimental statistics, it is found that when the image clarity, angle and illumination and other external adjustments are close to the ideal state, the detection results under the above similarity threshold selection have a test accuracy of 98.72% for handwritten text, a test accuracy of 99.96% for Simplified Chinese Songti text, and a test accuracy of 99.45% for Traditional Chinese Songti text.

[0088] In another optional embodiment of the present application, the second threshold value for judging the text sampling result is calculated according to the following formula in step S42:

[0089] calculating a second threshold value for judging the text sampling result;

[0090] wherein, a is the second threshold value for judging the text sampling result, τ is the definition of the image, τ' is the definition of the image in the ideal state, μ is the complexity of the image, the value range of μ is 0.5-1, and the lower the complexity of the image, the larger the value of μ, ρ is a constant parameter, and the size of ρ is determined according to the font type.

[0091] In this embodiment, by taking the ratio of the definition of the image to the definition of the image in the ideal state as the first parameter for calculating the second threshold value for judging the text sampling result, by taking the transformed result of the complexity of the image as the second parameter for calculating the second threshold value for judging the text sampling result, and by selecting different reference threshold values according to the font type, the obtained two parameters are multiplied by the reference threshold value to obtain the final second threshold value, which realizes the automatic and efficient monitoring and evaluation of the text quality, and provides strong support for the optimization and decision of the text processing process.

[0092] In the embodiment of the present application, after the image features are extracted and classified by using the clustering algorithm, and the text image is separated by using the morphological operation, at least two feature-separated text image data are obtained, including the following steps:

[0093] S51: high-level feature extraction, high-level feature extraction is performed on the at least one segmentation region, wherein the high-level feature extraction includes local binary pattern and gray level co-occurrence matrix;

[0094] S52: setting the clustering number K=n+2, wherein n is the number of times of detecting the segmentation region, and inputting the extracted high-level features and the comprehensive text feature data into the K-means clustering algorithm, and using the K-means clustering algorithm to divide the pixels and pixel blocks into n+2 clusters;

[0095] S53: removing the cluster represented by the background image, mapping the remaining clusters to the segmentation region, and separating the text region and removing the background noise by using the morphological operation, to obtain at least two feature-separated text image data.

[0096] Through the steps S111-S114, when recognizing the overlapped characters, the image is processed by using the clustering algorithm, the clustering algorithm is used for preprocessing and auxiliary separation, the text recognition algorithm is used for text recognition, the qualified detection is performed on the recognition result, and the clustering number of the clustering algorithm is changed according to the detection result, that is, the number of the recognized categories is sequentially overlapped from 3, and the recognition is performed again after the change, so that the accuracy of the recognition result can be effectively ensured, and the problems of low recognition success rate and accuracy of the overlapped characters in the existing text recognition process are solved.

[0097] In the embodiment of the present application, when the image post-processing is performed after the segmentation in the step S12 to further optimize and adjust the segmentation result, so as to improve the accuracy and stability of the recognition, the following steps are included:

[0098] S121: The erosion, dilation, opening operation, closing operation, morphological gradient operation, top-hat transformation and bottom-hat transformation in the morphological operation are used to perform the smoothing processing on the segmentation result, so as to remove the noise and burrs in the image;

[0099] S122: The repeated characters or broken characters in the segmentation result are removed and merged, so as to ensure that each character is independent and complete.

[0100] In another optional embodiment of the present application, the segmented image can also be smoothed according to the coordinates of the pixel points in the step S121, and the principle formula is as follows:

[0101]

[0102] wherein, I filtered (x,y) is the pixel value of the segmented image at the coordinates (x,y), sigma is the standard deviation of the Gaussian function, k is the size of the filter, i and j are loop variables for traversing all elements of the filter, and I(x+i,y+j) is the pixel value of the input image at the coordinates (x+i,y+j), that is, an element in the filter.

[0103] In this embodiment, the sum of the products of the k 2 pixels around each pixel point coordinate in the segmented image obtained and the Gaussian function is calculated, and then the sum of the product terms is divided to obtain the smoothed pixel value, so that the noise interference in the analysis of the segmented image can be suppressed, and the image is clearer.

[0104] In another optional embodiment of the present application, in the step S11, when the image containing the question content is acquired and segmented to obtain at least one segmentation region, the following steps are further included:

[0105] S111: Image enhancement, by adjusting the brightness, contrast, saturation, etc. of the image, enhancing the text information in the image, making the text more clear and identifiable;

[0106] S112: Image denoising, removing noise such as salt and pepper noise, Gaussian noise, etc. in the image, improving the clarity of the image;

[0107] S113: Image tilt correction, correcting the tilt angle of the image, keeping the text horizontal;

[0108] S114: Image grayscale processing, converting the image to a grayscale image for subsequent text detection and segmentation.

[0109] In this embodiment, by the above steps S111-S114, the image obtained by shooting is processed, the image quality problems caused by shooting technology, environment and shooting device clarity are weakened and eliminated, thereby improving the image quality and making the subsequent text detection and recognition more accurate.

[0110] In the embodiment of the application, when the text region is separated and the background noise is removed by morphological operation to obtain at least two feature-separated text image data, the following methods are included:

[0111] A11: Reduce the highlight area in the image by erosion operation, remove isolated noise points, and separate touching and approaching objects;

[0112] A12: Expand the highlight area in the image by dilation operation, fill the holes inside the object, and make the object edge smoother;

[0113] A13: Remove small objects by opening operation, separate objects at fine points, and smooth the boundaries of larger objects;

[0114] A14: Fill small holes inside the foreground object by closing operation, connect adjacent objects, and smooth their boundaries;

[0115] A15: Highlight the boundaries of objects by morphological gradient operation.

[0116] In the embodiment of the application, after text recognition on the at least two feature-separated text image data, the following steps are included:

[0117] S61: Position-based judgment, analyzing the position of the text in the image, matching it with the expected category or area;

[0118] S62: Check if the text content is consistent with the image context;

[0119] S63: Analyze the semantic meaning of the text, and determine whether it is consistent with the overall theme and context of the image text;

[0120] S64: Extract the original image of the recognized error character separately, and determine the possible meaning of the error character according to the semantic meaning of the text, and use the possible meaning of the error character as the judgment standard to re-identify the original image of the error character.

[0121] Through the above steps S61-S64, the technical scheme further enhances the analysis ability of the matching degree of the text and the image content on the basis of text recognition. First, based on the position judgment, the recognized text is accurately matched with the expected category or region in the image, which improves the accuracy of text positioning. Secondly, it checks whether the text content is consistent with the image context, ensuring the logical coherence of the text information. Thirdly, it deeply analyzes the semantic meaning of the text, and determines whether it is consistent with the overall theme and context of the image text, which further improves the accuracy and relevance of text recognition. Finally, for the recognized error character, the original image is extracted separately, and the possible meaning of the error character is inferred by combining the semantic meaning of the text, and the error character is re-identified as a standard, effectively reducing the overall text understanding deviation caused by single character recognition error. Not only improves the accuracy and efficiency of text recognition, but also enhances the relevance understanding of text and image content, making the text information closer to the actual situation of the image. In addition, through the intelligent re-identification mechanism of the error character, the text recognition error rate is further reduced, and the quality and user experience of the overall text processing are improved.

[0122] In the embodiment of the application, when at least one text content is selected for recognition and analysis according to the recognition result of the voice instruction data, the following steps are included:

[0123] S71: Voice signal conversion and preprocessing, converting the collected voice signal into a digital signal, and filtering, pre-emphasizing, framing and windowing the obtained digital signal;

[0124] S72: Speech feature extraction, converting the preprocessed digital signal from time domain to frequency domain, and extracting parameters that can represent the essential features of the speech signal;

[0125] S73: Acoustic modeling, using a hidden Markov model to match the extracted feature parameters with a preset acoustic model;

[0126] S74: Language modeling and decoding, modeling language characteristics through a recurrent neural network language model, and then using a decoder to fuse the information of the acoustic model and the language model to parse out the most possible text sequence.

[0127] In an embodiment of the present invention, when selecting at least one text content for recognition and analysis based on the recognition result of the voice instruction data, the following steps are also included:

[0128] S81: Text preprocessing, preprocessing the parsed text sequence, wherein the preprocessing methods used include noise removal, word segmentation and part-of-speech tagging;

[0129] S82: Named Entity Recognition, based on text preprocessing, performs named entity recognition to identify specific entities in the instructions;

[0130] S83: Intent recognition, using natural language processing technology to analyze the semantic content of instructions and identify the user's intention;

[0131] S84: Parameter extraction: based on identifying the user's intention, further extracting key parameters in the instruction;

[0132] S85: Instruction parsing and construction: Based on the user's intention and the extracted parameters, an instruction parsing tree or similar data structure is constructed to decompose complex instructions into a series of executable subtasks;

[0133] S86: Execute instructions.

[0134] In another optional embodiment of the present invention, when selecting at least one text content for recognition and analysis based on the recognition results of voice command data, the user may need to convert a large number of paper documents or scans into editable electronic documents, and perform content analysis and extraction. Through the embodiment of instruction parsing and construction, the user can input complex instructions, such as "identify all tables in the document and extract the data therein", and the system will automatically parse the instructions, decompose them into multiple sub-tasks such as text recognition, table positioning, data extraction, etc., and execute them in sequence, and finally output a structured data table.

[0135] In an embodiment of the present invention, when providing auxiliary guidance for a user's learning in at least one of voice, text, video, and image forms based on the recognition and analysis results of the text content, the following steps are included:

[0136] S91: Perform corresponding operations and responses according to the user's instructions;

[0137] S92: Use knowledge graph technology to deeply analyze the knowledge points contained in the question content, identify and extract key knowledge points, and use machine learning algorithms to predict the difficulties that users may encounter based on their learning history and performance;

[0138] S93: According to the obtained knowledge points and difficult points, search in the pre-set database, search for related audio, pictures and videos, and sort the searched audio, pictures and videos according to the set importance coefficient and the proportion coefficient of the knowledge points in the question, and take the order as the default playing order to play for the user.

[0139] In another optional embodiment of the application, the user can acquire images and watch videos and graphics through the AR device. When the user acquires images and watches videos and graphics through the AR device, the user first controls the AR device or other image device through voice instructions to collect the required images, and then performs text recognition through the above-mentioned image recognition steps. After the text recognition is completed, the user selects the required text content through voice instructions, and performs operations such as modification, arrangement and format conversion on the text content through voice instructions. After the operations are completed, the text content is parsed, and the knowledge points contained in the question content are analyzed in depth using the knowledge graph technology. Key knowledge points are identified and extracted, and the user's learning history and performance are combined to predict the difficulties that the user may encounter. Next, search is performed in the pre-set database to search for related tutoring materials, explanation videos, demonstration animations and related questions. When the question is explained and analyzed for the user, the searched related tutoring materials, explanation videos, demonstration animations and related questions are played through the AR device to deepen learning, ensuring the effectiveness of the user's learning.

[0140] The above is only the preferred embodiment of the application and is not used to limit the application. Although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art can modify the technical solutions recorded in the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.

Claims

1. AR-AI integrated interactive learning assistance method; characterized by: The following steps are included: S11: Acquire an image containing the topic content and segment the image to obtain at least one segmented region; S12: Recognize text patterns in at least one segmented area using image recognition technology to obtain text pattern data in at least one segmented area; S13: Extracting features of text patterns with different characteristics using feature recognition and classification technologies to obtain at least one text content; S14: Acquire user's voice command information data, and recognize the user's voice command data to obtain voice control instructions; S15: Select at least one text content for recognition and analysis based on the recognition result of the voice command data; S16: Based on the recognition and analysis results of the text content, provide auxiliary guidance for the user's learning in at least one of voice, text, video and image forms.

2. The AR-AI integrated interactive learning assistance method according to claim 1, characterized in that: When identifying text patterns in at least one segmented area by image recognition technology and obtaining text pattern data in the segmented area, the following steps are included: S21: Image preprocessing, enhancing text information in the image to make the text clearer and more legible; S22: Feature extraction: extract the features of the text in the image through edge detection technology and texture analysis technology to obtain comprehensive text feature data; S23: Image segmentation: by analyzing the connected domains in the image, the outline of the text is found and segmented according to the texture features; S24: Image post-processing, further optimization and adjustment of the segmentation results to improve the accuracy and stability of recognition.

3. The AR-AI integrated interactive learning assistance method according to claim 2, characterized in that: When extracting text patterns with different features separately through feature recognition technology and classification technology to obtain at least one text content, the following steps are included: S31: Text recognition: Based on the obtained comprehensive text feature data, the text is recognized to obtain a first preset text; S32: Performing a random inspection on the first preset text, inspecting 20%-40% of the characters in the first preset text; S33: When the sampling inspection result is qualified, the text recognition is ended; when the sampling inspection result is unqualified, the next step is entered; S34: First, the text patterns in at least one segmented area are separated using a color-based separation technology. If the separation is successful, at least two color-separated text image data are obtained. Text recognition is performed on the at least two separated text image data, and then the text is tested in step S32. If the test is qualified, the text recognition is terminated, otherwise the next step is entered; S35: Use a clustering algorithm to extract and classify image features, and use morphological operations to separate the text image to obtain at least two feature-separated text image data, and perform text recognition on at least two feature-separated text image data, and then perform inspection through step S32. After passing the inspection, the text recognition is terminated. Otherwise, the clustering parameters of the clustering algorithm are modified and recognition is performed again.

4. The AR-AI integrated interactive learning assistance method according to claim 3, characterized in that: When randomly sampling the first preset text and sampling 20%-40% of the characters in the first preset text, the following steps are included: S41: selecting a reference text, identifying the text texture data in the segmented area using optical character recognition technology, identifying the text font corresponding to the text texture data, and selecting a corresponding reference text according to the text content in the first preset text; S42: selecting a judgment threshold, selecting different similarity thresholds as a first threshold for judging character correctness based on the type of the identified text font, and calculating a second threshold for judging the text spot check result based on image clarity, font type, and image background complexity; S42: similarity calculation: randomly selecting characters in the first preset text, the number of characters selected being 20%-40% of all characters in the first preset text, and performing similarity calculation between images corresponding to the selected characters and images of the corresponding reference text; S43: Result judgment: judging whether the character is correct based on the calculated similarity value and the selected first threshold; S44: Comprehensive judgment: synthesize the judgment results of all the selected characters, calculate the total accuracy rate, and judge whether the sampling inspection result is qualified based on the obtained second threshold.

5. The AR-AI integrated interactive learning assistance method according to claim 4, characterized in that: The method comprises the following steps: extracting and classifying image features by using a clustering algorithm, and separating text images by using morphological operations to obtain at least two feature-separated text image data; and S51: Advanced feature extraction, performing advanced feature extraction on at least one segmented region, wherein the advanced feature extraction includes local binary pattern and gray level co-occurrence matrix; S52: Setting the number of clusters K=n+2, where n is the number of times the segmented area is detected, and inputting the extracted high-level features and comprehensive text feature data into a K-means clustering algorithm, using the K-means clustering algorithm to divide the pixels and pixel blocks into n+2 clusters; S53: Eliminate the clusters represented by the background image, map the remaining clusters to the segmented area, separate the text area and remove background noise through morphological operations, and obtain at least two feature-separated text image data.

6. The AR-AI integrated interactive learning assistance method according to claim 5, characterized in that: When separating the text area and removing background noise through morphological operations to obtain at least two feature-separated text image data, the following methods are included: A11: Uses erosion to reduce the highlight areas in the image, remove isolated noise points, and separate touching and approaching objects. A12: Dilation is used to expand the highlight areas in an image, fill holes inside objects, and smooth the edges of objects. A13: Removes small objects through opening operations, separates objects at thin points, and smoothes the boundaries of larger objects; A14: Fills small holes inside foreground objects through closing operations, connects adjacent objects, and smoothes their boundaries; A15: Highlighting object boundaries through morphological gradient operations.

7. The AR-AI integrated interactive learning assistance method according to claim 6, characterized in that: After performing text recognition on at least two feature-separated text image data, the method further includes the following steps: S61: Position-based judgment: analyzing the position of the text in the image and matching it with the expected category or area; S62: Check whether the text content is consistent with the image context; S63: Analyze the semantic meaning of the text to determine whether it fits the overall theme and context of the text in the image; S64: extracting the original image of the recognized incorrect character separately, judging the possible meaning of the incorrect character according to the semantic meaning of the text, and re-recognizing the original image of the incorrect character using the possible meaning of the incorrect character as a judgment criterion.

8. The AR-AI integrated interactive learning assistance method according to claim 7, characterized in that: The steps of obtaining the user's voice command information data, recognizing the user's voice command data, and obtaining the voice control command include: S71: Voice signal conversion and preprocessing, converting the collected voice signal into a digital signal, and performing filtering, pre-emphasis, framing, and windowing on the obtained digital signal; S72: Speech feature extraction, converting the preprocessed digital signal from the time domain to the frequency domain, and extracting parameters that can characterize the essential characteristics of the speech signal; S73: Acoustic modeling, using a hidden Markov model to match the extracted feature parameters with the preset acoustic model; S74: Language modeling and decoding, modeling language characteristics through a recurrent neural network language model, and then using a decoder to fuse the information of the acoustic model and language model to parse the most likely text sequence.

9. The AR-AI integrated interactive learning assistance method according to claim 8, characterized in that: When selecting at least one text content for recognition and analysis based on the recognition result of the voice command data, the following steps are included: S81: Text preprocessing, preprocessing the parsed text sequence, wherein the preprocessing methods used include noise removal, word segmentation and part-of-speech tagging; S82: Named Entity Recognition, which performs named entity recognition based on text preprocessing to identify specific entities in the instructions; S83: Intent recognition, using natural language processing technology to analyze the semantic content of instructions and identify the user's intention; S84: Parameter extraction: based on identifying the user's intention, further extracting key parameters in the instruction; S85: Instruction parsing and construction: Based on the user's intention and the extracted parameters, an instruction parsing tree or similar data structure is constructed to decompose complex instructions into a series of executable subtasks; S86: Execute instructions.

10. The AR-AI integrated interactive learning assistance method according to claim 9, characterized in that: When providing auxiliary guidance for the user's learning in at least one of voice, text, video and image forms based on the recognition and analysis results of the text content, the following steps are included: S91: Perform corresponding operations and responses according to the user's instructions; S92: Use knowledge graph technology to deeply analyze the knowledge points contained in the question content, identify and extract key knowledge points, and use machine learning algorithms to predict the difficulties that users may encounter based on their learning history and performance; S93: Based on the obtained knowledge points and difficulty points, search for relevant audio, pictures and videos in a pre-set database, and sort the searched audio, pictures and videos according to the set importance coefficient and the proportion coefficient of the knowledge point in the question, and use this order as the default playback order to play for the user.

Citation Information

Patent Citations

  • Text layout analysis method and device, computer equipment and storage medium

    CN111340037A

  • Image character recognition method and system based on deep learning and medium

    CN112016547A

  • Text information recognition method and device

    CN112906499A

  • Label element machine vision quality inspection method and device

    CN115983302A

  • System and method for assisting in learning foreign language based on large model, and user terminal

    CN118135855A