Braille auxiliary learning system and method

By designing a Braille assisted learning system that includes image acquisition, image processing, information conversion and central processing modules, combined with gesture recognition and text analysis technology, the voice broadcast or Braille display of text is realized by blind users through natural gesture movements, solving the problems of single functions and cumbersome operations in the existing system, significantly improving learning efficiency and Braille accuracy.

CN120014915APending Publication Date: 2025-05-16HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510447704.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing Braille assisted learning system has a single function and lacks efficient Chinese blind conversion display function. It is cumbersome and expensive to operate, making it difficult to lower the threshold for blind people to learn and improve learning efficiency.

Method used

A Braille assisted learning system is designed, including an image acquisition module, an image processing module, an information conversion module and a central processing module. The camera collects environmental images in real time, and combines gesture recognition and text analysis technology to realize real-time interpretation of user gestures. Users can control the voice broadcast or Braille display of text through natural gesture actions.

Benefits of technology

Through gesture control function, the system translates Chinese characters into Braille or converts them into natural speech in real time, which significantly improves the operation convenience and learning efficiency of blind users, effectively avoids reading difficulties that may be caused by improper tone omission in current Braille, and improves the accuracy and readability of Braille.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014915A_ABST
    Figure CN120014915A_ABST
Patent Text Reader

Abstract

The invention discloses a braille auxiliary learning system and method, the system comprises an image acquisition module, an image processing module, an information conversion module and a central processing module, the image acquisition module is used for acquiring an environment image, the image processing module is used for analyzing and processing the environment image to acquire image information, and the image information is transmitted to the central processing module; the information conversion module comprises a voice broadcast unit and a braille display unit, the voice broadcast unit is used for converting text content into audio data for voice broadcast, the braille display unit is used for converting the text content into braille data for braille display, and the central processing module is used for displaying the braille according to the type of the gesture instruction. And selectively driving the voice broadcast unit to perform voice broadcast on the text content, or controlling the braille display unit to perform braille display on the text content. According to the system, the time cost and the economic cost in the braille learning process are reduced, and meanwhile, the braille learning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Braille learning, and in particular to a Braille auxiliary learning system and method. Background Art

[0002] Braille books are bulky and lack variety, which has long limited the reading options of the visually impaired. Compared with ordinary books, Braille books are expensive to produce, and the conversion process from paper books to Braille books requires software scanning, manual proofreading, and professional machine production, which is not only time-consuming and labor-intensive, but also prone to errors. In addition, there are few professional Braille publishing houses and the market competition is not fierce, resulting in insufficient innovation motivation for publishing houses, and limited room for improvement in the content and quality of Braille books. In the process of using Braille books, there is also the problem that books are easily worn out due to repeated use and become smooth and unusable, and new books may scratch fingers due to sharp edges.

[0003] In the existing technology, solutions for blind people to learn Braille generally have problems such as high cost, low level of technical research and development, and limited promotion channels. For example, the existing auxiliary Braille learning devices are limited in number, single in function, lack efficient Chinese-Blind conversion display function, and are cumbersome to operate, expensive, and bulky. Therefore, it is urgent to develop an intelligent Braille auxiliary learning system that can lower the learning threshold and improve learning efficiency. Summary of the invention

[0004] The purpose of the present invention is to provide a Braille-assisted learning system and method, which, through the gesture control function, can translate Chinese characters into Braille in real time according to learning needs and provide users with touch learning through a Braille display, or convert Chinese characters into natural voice for broadcasting, thereby solving the problems of single function and lack of efficient Chinese-Braille conversion in the prior art Braille-assisted learning system.

[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0006] The present invention provides a Braille auxiliary learning system, which comprises: an image acquisition module, an image processing module, an information conversion module and a central processing module;

[0007] The image acquisition module is used to acquire an environment image;

[0008] The image processing module is used to analyze and process the environment image to obtain image information, wherein the image information includes text content or gesture instructions;

[0009] The information conversion module includes a voice broadcast unit and a braille display unit, wherein the voice broadcast unit is used to convert the text content into audio data for voice broadcast, and the braille display unit is used to convert the text content into braille data for braille display;

[0010] The central processing module is used to selectively drive the voice broadcast unit to voice broadcast the text content, or control the braille display unit to display the text content in braille according to the type of the gesture instruction.

[0011] In one embodiment of the present invention, the image acquisition module includes a camera, and the image acquisition module collects an environmental image through the camera and sends the environmental image to the image processing module.

[0012] In one embodiment of the present invention, the image processing module includes an image recognition unit, a text parsing unit and a gesture recognition unit. The image recognition unit is used to distinguish the image type of the environmental image according to the image features of the environmental image. If the environmental image is a text image, the text image is sent to the text parsing unit. If the environmental image is a gesture image, the gesture image is sent to the gesture recognition unit.

[0013] In one embodiment of the present invention, the text parsing unit is used to extract text content from the text image and send the text content to the central processing module, and the gesture recognition unit is used to parse the gesture image to determine the corresponding gesture instructions and send the gesture instructions to the central processing module.

[0014] In one embodiment of the present invention, the text parsing unit extracts text content from the text image using optical character recognition technology, and sends the extracted text content to the central processing module.

[0015] In one embodiment of the present invention, the gesture recognition unit uses a preset model to parse the gesture image to extract gesture information, and then compares and matches the extracted gesture information with a preset gesture instruction set to obtain corresponding gesture instructions, and sends the gesture instructions to the central processing module, wherein the gesture information includes static gesture information and dynamic gesture information.

[0016] In one embodiment of the present invention, the Braille display unit includes a Chinese-Braille encoding subunit and a Braille dot display, and the Chinese-Braille encoding subunit is used to convert the text content into Braille data and send the Braille data to the Braille dot display for Braille display.

[0017] In one embodiment of the present invention, the braille display comprises an electromagnetic braille display and a cam piston braille display.

[0018] In one embodiment of the present invention, the system also includes a mobile interaction module, which is communicatively connected to the central processing module. The mobile interaction module is used to parse the text image taken by the mobile terminal to obtain structured text data, and send the structured text data to the central processing module. The central processing module controls the Braille display unit to convert the structured text data into Braille data for Braille display.

[0019] Based on the same inventive concept, another embodiment of the present invention further provides a Braille assisted learning method, which adopts the Braille assisted learning system as described in any of the above embodiments, and the method includes:

[0020] Acquire environmental images through a camera;

[0021] Analyzing and processing the environment image by using an image processing unit to extract image information from the environment image, wherein the image information includes text content and gesture instructions;

[0022] According to the type of the gesture instruction, the voice broadcast unit is selectively driven to convert the extracted text content into voice data for voice broadcast, or the braille display unit is controlled to convert the extracted text content into braille data for braille display.

[0023] As described above, the present invention provides a Braille assisted learning system, including an image acquisition module, an image processing module, an information conversion module and a central processing module, wherein the image acquisition module is used to acquire an environmental image, the image processing module is used to analyze and process the environmental image to acquire image information, wherein the image information includes text content or gesture instructions, the information conversion module includes a voice broadcast unit and a Braille display unit, the voice broadcast unit is used to convert the text content into audio data for voice broadcast, the Braille display unit is used to convert the text content into Braille data for Braille display, and the central processing module is used to selectively drive the voice broadcast unit to voice broadcast the text content, or control the Braille display unit to display the text content in Braille according to the type of the gesture instruction. The system collects gesture images and reading material images in real time through a camera, and combines gesture recognition and text parsing technology to realize real-time interpretation of user gestures. Users only need to use natural gestures to easily control the voice broadcast or Braille display of text, without complicated key operations, thereby greatly improving the convenience of operation for blind users. At the same time, the gesture recognition module of the system not only supports the recognition of static gestures, but also can accurately capture dynamic gesture movements, such as pinching and grasping of the hands, etc. These functional designs are in line with the natural habits of the blind during the reading process. In addition, the system is equipped with an efficient voice broadcast function, which improves the user experience from both the action and hearing aspects. In terms of Chinese-Blind conversion, the system can accurately implement the tone omitting and abbreviation rules, and the accuracy of Chinese-Blind conversion is high, which fully meets the actual needs of blind people in learning and reading, and effectively avoids the reading difficulties that may be caused by improper omission of tones in the current Braille, and significantly improves the accuracy and readability of Braille. At the same time, the system performs well in processing long corpus files, and the processing speed far exceeds the blind people's touch reading speed. It is suitable for real-time interactive scenarios and the system has high execution efficiency. Of course, it is not necessary to achieve all the advantages described above at the same time for any product implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 A structural schematic diagram of a Braille assisted learning system provided as an exemplary embodiment of the present application.

[0026] Figure 2 A hardware-side workflow diagram of a Braille-assisted learning system provided as an exemplary embodiment of the present application.

[0027] Figure 3 A schematic structural diagram of a Braille display provided as an exemplary embodiment of the present application.

[0028] Figure 4 A mobile interactive end workflow diagram of a Braille assisted learning system provided as an exemplary embodiment of the present application.

[0029] Figure 5 A flowchart of a Braille assisted learning method provided by an exemplary embodiment of the present application.

[0030] The reference numerals are as follows:

[0031] 100 Image acquisition module

[0032] 200 Image processing module

[0033] 210 Image recognition unit

[0034] 220 Text parsing unit

[0035] 230 Gesture Recognition Unit

[0036] 300 Central Processing Module

[0037] 400 Information conversion module

[0038] 410 Braille display unit

[0039] 420 Voice Broadcast Unit

[0040] 1 Tablet

[0041] 2 dot display unit

[0042] 3. First base

[0043] 4 First hinge

[0044] 5 First connecting rod

[0045] 6 Second base

[0046] 7 Second hinge

[0047] 8 Second connecting rod DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0049] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0050] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, structures and devices known to the public are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0051] In order to solve the technical problems that the existing Braille-assisted learning system has a single function and lacks efficient Chinese-Braille conversion, the present invention provides a Braille-assisted learning system. The system has a gesture control function, can respond to learning needs in real time, and choose to translate Chinese characters into Braille and display them through a Braille display for users to touch and learn according to gesture instructions, or convert Chinese characters into natural voice for broadcasting, thereby reducing the time and economic costs in the Braille learning process. At the same time, the system is also equipped with a camera that can capture and recognize text information in real time, thereby providing personalized learning content for blind users and further enriching learning resources.

[0052] See also Figure 1As shown, in an exemplary embodiment of the present application, the braille assisted learning system includes an image acquisition module 100, an image processing module 200, an information conversion module 400 and a central processing module 300. The image acquisition module 100 is used to acquire an environmental image; the image processing module 200 is used to analyze and process the environmental image to acquire image information, wherein the image information includes text content or gesture instructions; the information conversion module 400 includes a voice broadcast unit 410 and a braille display unit 420, the voice broadcast unit 410 is used to convert the text content into audio data for voice broadcast, and the braille display unit 420 is used to convert the text content into braille data for braille display; the central processing module 300 is used to selectively drive the voice broadcast unit 410 to voice broadcast the text content, or control the braille display unit 420 to display the text content in braille according to the type of the gesture instruction.

[0053] It should be noted that the image acquisition module 100 is connected to the image processing module 200, and the image acquisition module 100 includes a camera and an image sensor. In the present embodiment, the image acquisition module 100 uses a camera. The image acquisition module 100 collects the environmental image around the system equipment through the camera and sends the environmental image to the image processing module 200 for subsequent processing.

[0054] Please continue reading Figure 1 As shown, in an exemplary embodiment of the present application, the image processing module 200 includes an image recognition unit 210, a text parsing unit 220 and a gesture recognition unit 230. The image recognition unit 210 is used to distinguish the image type of the environmental image according to the image features of the environmental image. If the environmental image is a text image, the text image is sent to the text parsing unit 220. If the environmental image is a gesture image, the gesture image is sent to the gesture recognition unit 230.

[0055] Specifically, the image recognition unit 210 is connected to the image acquisition module 100, the text parsing unit 220 and the gesture recognition unit 230 respectively, and is used to analyze the environmental image captured by the camera, and judge the type of the environmental image based on image features (such as edge density, shape complexity, etc.) to distinguish whether the environmental image is a text image or a gesture image. In this embodiment, the text image contains flat text such as books and screens, and the gesture image contains user hand movements. If the environmental image is judged to be a text image, the environmental image is sent to the text parsing module 220 for further processing, and if the environmental image is judged to be a gesture image, the environmental image is sent to the gesture recognition module 230 for further processing.

[0056] In an exemplary embodiment of the present application, the text parsing unit 220 is used to extract text content from the text image and send the text content to the central processing module 300, and the gesture recognition unit 230 is used to parse the gesture image to determine the corresponding gesture instruction and send the gesture instruction to the central processing module 300.

[0057] It should be noted that, in this embodiment, the text parsing unit 220 extracts the text content in the text image using optical character recognition (OCR) technology, and sends the extracted text content to the central processing module 300 .

[0058] Specifically, the text parsing unit 220 includes a grayscale processing subunit, a binarization subunit, an edge detection subunit, a document correction subunit, a sharpening subunit and an OCR recognition subunit. The grayscale processing subunit, the binarization subunit and the edge detection subunit are connected in sequence, the document correction subunit is respectively connected to the binarization subunit and the edge detection subunit, and the sharpening subunit is respectively connected to the document correction subunit and the OCR recognition subunit. It should be noted that, in this embodiment, the sharpening subunit adopts a neural network, and its network parameters can be obtained through adversarial training. The training of the sharpening subunit specifically includes the following steps:

[0059] SA1: By obtaining a basic model composed of a grayscale processing subunit, a binarization subunit, an edge detection subunit, a document correction subunit, an encoding subunit, a decoding subunit and an OCR recognition subunit, the text image is converted into a grayscale image after grayscale processing by the grayscale processing subunit, then the grayscale image is converted into a binary image by the binarization subunit, then the edge detection subunit detects the contour of the binary image, the document correction subunit corrects the contour of the binary image, and then obtains the corrected binary image, the encoding subunit and the decoding subunit both use neural networks, the encoding subunit sharpens and encodes the corrected binary image, and inputs the obtained sharpened image into the OCR recognition subunit for recognition to obtain the text recognition result, and the decoding subunit decodes the sharpened image output by the encoding subunit to output a decoded image.

[0060] SA2: Construct a learning sample, which is a reading material image annotated with real text information, and construct a loss function, which is the sum of a first loss and a second loss. The first loss is the difference between the corrected binary image input to the encoding subunit and the decoded image output by the decoding subunit, and the second loss is the difference between the text recognition result output by the OCR recognition subunit and the real text information corresponding to the reading material image.

[0061] SA3: The basic model learns the learning samples and updates the encoding subunit and decoding subunit through the back propagation of the loss function. When the basic model converges, the encoding subunit is extracted as the sharpening subunit.

[0062] The OCR recognition subunit is used to recognize and convert the text information in the image into a text format that can be edited and processed by a computer. The text detection method based on deep learning, such as convolutional neural network (CNN) or region proposal network (RPN), can automatically locate and segment the text area in the image, and then use algorithms such as recurrent neural network (RNN) or long short-term memory network (LSTM) to perform character recognition and conversion on the detected text area, and convert the text in the image into a computer-readable text format.

[0063] In an exemplary embodiment of the present application, the gesture recognition unit 230 uses a preset model to parse the gesture image to extract gesture information, and then compares and matches the extracted gesture information with a preset gesture instruction set to obtain corresponding gesture instructions, and sends the gesture instructions to the central processing module 300, wherein the gesture information includes static gesture information and dynamic gesture information.

[0064] It should be noted that, in this embodiment, the gesture recognition module 230 matches the acquired gesture information through the central processing module 300, specifically including data collection and processing, model construction and training, model evaluation and prediction, and gesture information analysis, so as to realize real-time gesture recognition, and control the voice broadcast unit 410 or the Braille display unit 420 according to the matching gesture instructions.

[0065] Specifically, the data collection and processing is to capture video frames through a camera and detect hand feature points in the video frames. In this embodiment, the data collection process includes 13 gestures, namely "one", "two", "three", "four", "five", "six", "seven", "eight", "nine", "ten", "good", "not good", and "ok". 300 data points are collected for each gesture, and each data point contains the spatial coordinates (x, y, z) of 21 hand feature points. The hand in the video frame is identified and located by the detecHands method in the HandDetector class, and the three-dimensional coordinates of the 21 key feature points of the hand are extracted. The detected hand feature point coordinates are stored as a Numpy array, and the length of each data point is 21 feature points * 3 coordinate dimensions = 63.

[0066] The model construction and training, that is, building a neural network model, including several fully connected layers and ReLU activation functions, and finally using the softmax activation function for classification. The data set is loaded through the loadData method, and the loaded data set is divided into a training set and a test set, with a division ratio of 80:20, that is, 80% of the data is used to train the model, and the remaining 20% ​​of the data is used to test the performance of the model. The model is compiled using the Adam optimizer and the sparse classification cross entropy loss function, and trained for 1000 epochs. Epoch refers to the process in which the entire data set passes through the model once (forward propagation and back propagation). Through multiple iterative training, the model can gradually learn the features in the data and optimize its parameters to minimize the loss function. After the training is completed, the model can recognize and classify gesture information, and store the mapping relationship between the gesture information obtained from the training and the gesture instructions. The gesture instructions include Braille display, page turning, volume adjustment (the user can map 0-100 volume control according to the distance between two fingers), shutdown, etc.

[0067] The model evaluation and prediction, that is, during the test process, real-time video frames are captured by a camera, the trained model is used for prediction, and the performance of the model is evaluated based on the accuracy of the prediction results.

[0068] The gesture information analysis, that is, the gesture recognition module 230 extracts gesture information from the gesture image and matches it with the preset gesture instruction set one by one to determine the gesture instruction corresponding to the gesture image. The gesture information includes static gesture information and dynamic gesture information.

[0069] In an exemplary embodiment of the present application, the central processing module 300 uses a Raspberry Pi main control module to facilitate the deployment of various functional modules and improve the system computing efficiency and performance. The central processing module 300 obtains gesture information and text content, and then matches the acquired gesture information with a preset gesture instruction set. According to the matched gesture instruction, the central processing module 300 selectively drives the voice broadcast unit 410 to convert the text content into audio data for playback, or controls the braille display unit 420 to convert the text content into braille data for display.

[0070] In an exemplary embodiment of the present application, the braille display unit 420 includes a Chinese-Blind encoding subunit and a braille dot display, and the Chinese-Blind encoding subunit is used to convert the text content into braille data, and send the braille data to the braille dot display for braille display. It should be noted that the braille dot display includes an electromagnetic braille dot display and a cam piston braille dot display, and the braille data is six-bit binary data.

[0071] It is worth noting that in this embodiment, please refer to Figure 3As shown, the braille display adopts an electromagnetic braille display, which is composed of eight independent display units 2, each of which is equipped with six display devices, and the vertical displacement of the display devices is independently controlled by telescopic electromagnets. When the electromagnet is not powered, the top of the contact is horizontally aligned with the flat plate 1. After power is turned on, the electromagnet pushes the contact to move upward so that the top of the contact is higher than the flat plate plane. By arranging these six contacts in an orderly manner, a braille syllable can be formed for visually impaired users to recognize.

[0072] Specifically, the top of the braille display is a flat plate 1, which is divided into eight independent display units 2. Each display unit 2 is provided with six through holes arranged in a 2×3 matrix, corresponding to the contact position. Six constraint sleeves are provided below each display unit 2, which cooperate with the extension arm below the contact to ensure that the extension arm can only move in the vertical direction. An extension arm is connected below the contact, and a sleeve is provided below the other end of the extension arm. The sleeve is sleeved on the end of the electromagnet output shaft and moves up and down with it. The telescopic electromagnet is composed of a rectangular body and an output shaft. The electromagnet body is fixedly mounted on the electromagnet base, and the end of the output shaft cooperates with the sleeve. The vertical movement of the six electromagnet output shafts is achieved through the integration of the sleeve-contact mechanism to realize the vertical displacement of the six contacts in one display unit 2. The eight display units 2 together constitute a complete braille display.

[0073] The box body is connected to the electromagnet base. A first base 3 of a first connecting rod 5 is provided on the outer side of the right side of the box body, and the first base 3 is connected to the first connecting rod 5 by a first hinge 4. The first hinge 4 allows the first connecting rod 5 to stably hover at any angle by an interference fit with the first connecting rod 5. The first connecting rod 5 can rotate around the x-axis. A second base 6 connected to a second connecting rod 8 is provided at the end of the first connecting rod 5, and the second connecting rod 8 cooperates with the first connecting rod 5 by a second hinge 7, and the second connecting rod 8 can rotate around the second hinge 7 perpendicular to the x-axis. The second connecting rod 8 also has an interference fit with the second hinge 7, allowing the second connecting rod 8 to stably hover at any angle. A camera (not shown) is installed at the end of the second connecting rod 8. In the unfolded state, the camera is aligned downward with the central axis of the box body. In the retracted state, the two connecting rods are folded on the right side of the box body and are in a horizontal state.

[0074] It should be noted that, in the present embodiment, the Chinese-Blind coding subunit is a Chinese-Blind conversion system based on the national general Braille tone marking rules, and the Chinese-Blind conversion system includes initialization data preparation, word segmentation function and conversion function. Specifically, the initialization data preparation, that is, according to the national general Braille tone marking rules, defines BlindTree1 and BlindTree2, which are respectively used for the Braille coding trees of initials and finals. Phonetic_symbols, dictionary and dictionaryB are defined for the Braille coding of tones, punctuation marks, letters and numbers. Node class and BlindTree1, BlindTree2 classes are defined to construct the tree structure of initials and finals. The word segmentation function, that is, uses the jieba library to perform word segmentation, segments the input string and returns the word segmentation result (i.e. the obtained pinyin). The conversion function, that is, converts the word segmentation result into Braille coding according to the defined Braille coding tree.

[0075] See also Figure 2 As shown, Figure 2 The flowchart of the hardware end of the Braille assisted learning system is shown. After the device is powered on, it automatically broadcasts voice to prompt the gestures corresponding to different functions. The visually impaired user unlocks the corresponding function through gestures. If the learning mode is selected, the system turns on the braille display function, displays Braille through the braille display device, and broadcasts the displayed text by voice to achieve the learning effect.

[0076] In an exemplary embodiment of the present application, the Braille assisted learning system also includes a mobile interaction module, which is communicatively connected to the central processing module 300. The mobile interaction module is used to parse the text image taken by the mobile terminal to obtain structured text data, and send the structured text data to the central processing module 300. The central processing module 300 controls the Braille display unit 420 to convert the structured text data into Braille data for Braille display.

[0077] Specifically, the mobile interaction module is connected to the central processing module 300 through a local area network, and realizes the functions of taking pictures and voice control by calling relevant APIs, and uses OCR recognition technology to parse the text image obtained by taking pictures, and then transmits the parsed structured text data to the central processing module 300. The central processing module 300 receives the structured text data and performs corresponding encoding conversion, and displays braille through a braille display, thereby realizing the mobile terminal expansion of the braille assisted learning system. Figure 4As shown, after the mobile system is started, the voice module is initialized first to ensure that the voice synthesis function is normal, and then the voice recognition function is started to prepare to receive the user's voice command. If the user selects the audio playback function, the local specified book will start to play. The system determines whether the book content is successfully extracted. If so, it starts to play the specified book content. If not, it prompts the user to reselect or confirm the book location through voice. If the user selects the camera function, the camera is triggered according to the user's instructions or automatically. The user aims at the book and triggers the photo operation. The system recognizes the text of the collected image and prepares for voice broadcast.

[0078] Based on the same inventive concept, please refer to Figure 5 As shown, the present invention also provides a Braille assisted learning method, which adopts the Braille assisted learning system as described in any of the above embodiments, and the method comprises the following steps:

[0079] S100: Acquire an environment image through a camera;

[0080] S200: Analyzing and processing the environment image by using an image processing unit to extract image information from the environment image, wherein the image information includes text content and gesture instructions;

[0081] S300: According to the type of the gesture instruction, selectively drive the voice broadcast unit to convert the extracted text content into voice data for voice broadcast, or control the Braille display unit to convert the extracted text content into Braille data for Braille display.

[0082] In summary, the present invention provides a braille assisted learning system, including an image acquisition module 100, an image processing module 200, an information conversion module 400 and a central processing module 300, wherein the image acquisition module 100 is used to acquire an environmental image, the image processing module 200 is used to analyze and process the environmental image to acquire image information, wherein the image information includes text content or gesture instructions, the information conversion module 400 includes a voice broadcast unit 410 and a braille display unit 420, the voice broadcast unit 410 is used to convert the text content into audio data for voice broadcast, the braille display unit 420 is used to convert the text content into braille data for braille display, and the central processing module 300 is used to selectively drive the voice broadcast unit 410 to voice broadcast the text content, or control the braille display unit 420 to display the text content in braille according to the type of the gesture instruction. The system collects gesture images and reading material images in real time through a camera, and combines gesture recognition and text parsing technology to realize real-time interpretation of user gestures. Users can easily control the voice broadcast or Braille display of text through natural gestures, without complicated key operations, which greatly improves the convenience of operation for blind users. At the same time, the gesture recognition module of the system not only supports the recognition of static gestures, but also can accurately capture dynamic gestures, such as pinching of the hands, etc. These functional designs are in line with the natural habits of the blind during reading. In addition, the system is equipped with an efficient voice broadcast function, which improves the user experience from both the action and hearing aspects. In terms of Chinese-Blind conversion, the system can accurately implement the rules of tone omission and abbreviation, and the accuracy of Chinese-Blind conversion is high, which fully meets the actual needs of blind people in learning and reading, effectively avoids the reading difficulties that may be caused by improper omission of tones in the current Braille, and significantly improves the accuracy and readability of Braille. At the same time, the system performs well in processing long corpus files, and the processing speed is far faster than the blind people's touch reading speed. It is suitable for real-time interactive scenarios and the system has high execution efficiency.

[0083] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A Braille assisted learning system, characterized in that: include: Image acquisition module, image processing module, information conversion module and central processing module; The image acquisition module is used to acquire an environment image; The image processing module is used to analyze and process the environment image to obtain image information, wherein the image information includes text content or gesture instructions; The information conversion module includes a voice broadcast unit and a braille display unit, wherein the voice broadcast unit is used to convert the text content into audio data for voice broadcast, and the braille display unit is used to convert the text content into braille data for braille display; The central processing module is used to selectively drive the voice broadcast unit to voice broadcast the text content, or control the braille display unit to display the text content in braille according to the type of the gesture instruction.

2. The Braille assisted learning system according to claim 1, characterized in that: The image acquisition module includes a camera, and the image acquisition module collects an environmental image through the camera and sends the environmental image to the image processing module.

3. The Braille assisted learning system according to claim 1, characterized in that: The image processing module includes an image recognition unit, a text parsing unit and a gesture recognition unit. The image recognition unit is used to distinguish the image type of the environmental image according to the image features of the environmental image. If the environmental image is a text image, the text image is sent to the text parsing unit. If the environmental image is a gesture image, the gesture image is sent to the gesture recognition unit.

4. The Braille assisted learning system according to claim 3, characterized in that: The text parsing unit is used to extract text content from the text image and send the text content to the central processing module. The gesture recognition unit is used to parse the gesture image to determine the corresponding gesture instruction and send the gesture instruction to the central processing module.

5. The Braille assisted learning system according to claim 4, characterized in that: The text parsing unit extracts text content from the text image using optical character recognition technology, and sends the extracted text content to the central processing module.

6. The Braille assisted learning system according to claim 4, characterized in that: The gesture recognition unit uses a preset model to analyze the gesture image to extract gesture information, then compares and matches the extracted gesture information with a preset gesture instruction set to obtain corresponding gesture instructions, and sends the gesture instructions to the central processing module, wherein the gesture information includes static gesture information and dynamic gesture information.

7. The Braille assisted learning system according to claim 1, characterized in that: The Braille display unit includes a Chinese-Braille encoding subunit and a Braille dot display. The Chinese-Braille encoding subunit is used to convert the text content into Braille data and send the Braille data to the Braille dot display for Braille display.

8. The Braille assisted learning system according to claim 7, characterized in that: The braille display device includes an electromagnetic braille display device and a cam piston braille display device.

9. The Braille assisted learning system according to claim 1, characterized in that: The system also includes a mobile interaction module, which is communicatively connected to the central processing module. The mobile interaction module is used to parse the text image taken by the mobile terminal to obtain structured text data, and send the structured text data to the central processing module. The central processing module controls the Braille display unit to convert the structured text data into Braille data for Braille display.

10. A Braille assisted learning method, characterized in that: The method adopts the Braille assisted learning system according to any one of claims 1 to 9, and the method comprises: Acquire environmental images through a camera; Analyzing and processing the environment image by using an image processing unit to extract image information from the environment image, wherein the image information includes text content and gesture instructions; According to the type of the gesture instruction, the voice broadcast unit is selectively driven to convert the extracted text content into voice data for voice broadcast, or the braille display unit is controlled to convert the extracted text content into braille data for braille display.