Multi-mode tactile and visual fusion recognition method for medicine box braille salient point characters

Through the multimodal tactile visual fusion recognition method, the problem of low accuracy and insufficient coordination of Braille recognition technology in complex environments is solved, high-precision Braille recognition is achieved, and multi-channel feedback is provided to support the safe use of drugs for visually impaired people.

CN120277608APending Publication Date: 2025-07-08NANJING UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510342650.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有的盲文识别技术在环境复杂和药盒材质多样情况下准确率低,多模态融合技术协同不足,无法满足视障人士安全用药的高精度需求。

Method used

The multimodal haptic visual fusion recognition method is adopted to synchronize data through the haptic sensor array and vision sensor, perform spatiotemporal alignment and feature-level decision-making fusion, combine machine learning models to perform semantic analysis of Braille characters, and provide recognition results through a multi-channel feedback module.

Benefits of technology

It significantly improves the accuracy and environmental adaptability of Braille recognition, and can accurately identify Braille medicine boxes in complex environments to meet the safe medication needs of visually impaired people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277608A_ABST
    Figure CN120277608A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode tactile and visual fusion identification method for medicine box braille salient point characters. The method comprises the following steps: acquiring three-dimensional morphology and pressure data of Braille salient points through a tactile sensor, acquiring an image and the like through a visual sensor, and performing preprocessing (improved U-Net segmentation, illumination correction and the like); hardware is utilized to trigger synchronization, an ICP algorithm is improved to complete space-time alignment, and a feature level (attention mechanism) and a decision level (D-S evidence theory) are fused; character semantics are analyzed by means of a multi-mode neural network, and multi-channel feedback is visualized through voice, tactile vibration and AR. A dynamic optimization mechanism is set, and sensor scheduling, incremental learning and safety verification are covered. The system comprises a tactile / visual perception module, a data processing module (FPGA + NPU), a feedback execution module and a communication module. According to the scheme, the problems of low traditional recognition precision and poor environmental adaptability are solved, high-precision medicine box braille recognition is achieved, and a reliable scheme suitable for multiple scenes is provided for safe medicine use of visually impaired people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and sensor fusion, and particularly to a multi-modal tactile-visual fusion recognition method for Braille bump characters on medicine boxes. Background Art

[0002] In daily life, visually impaired people rely on Braille to obtain information, and the Braille on medicine boxes is crucial for them to accurately identify medicines. However, current Braille recognition technologies have many defects, bringing great inconvenience to the medicine use of visually impaired people.

[0003] Traditional single-modal Braille recognition technologies have many problems. Taking tactile recognition as an example, the commonly used hand-held tactile recognition devices for visually impaired people can only sense the surface pressure of Braille bumps. This is like roughly touching with fingers, only knowing where there are bumps, but it is difficult to know the exact height and shape of the bumps. In the actual medicine box use scenario, the materials of medicine boxes are diverse. For some plastic medicine boxes, the Braille bumps become flat due to long-term use and wear, and for some paper medicine boxes, the Braille bumps are not clear enough originally. Facing these situations, the recognition accuracy of such devices is often lower than 80%, resulting in visually impaired people may misjudge medicines. And although visual recognition technology analyzes the images collected by cameras, in the real environment, the light conditions are complex and changeable. In a dim indoor environment, the Braille images collected by the camera are blurred; in the outdoor environment with direct sunlight, the surface of the medicine box reflects light severely, making the recognition system make mistakes frequently. Moreover, visual recognition cannot directly obtain the height information of Braille bumps, and it is difficult to accurately recognize some Braille with different heights due to the manufacturing process on the medicine box.

[0004] Although multi-modal fusion technology has developed, the existing solutions are not satisfactory. Some simple methods of splicing tactile and visual data are like simply bundling two tools together without a collaborative working mechanism. In actual use, when visually impaired people hold the device to recognize the Braille on the medicine box in different environments, such as on a bumpy bus, the tactile data and visual data cannot be effectively fused due to poor spatio-temporal alignment accuracy, and the recognition effect is greatly reduced. And there are various variants of Braille on medicine boxes, and the font sizes, spacings, bump heights, etc. of Braille on medicine boxes of different pharmaceutical factories are different. The existing technologies are difficult to adapt to these changes and cannot meet the urgent need of visually impaired people for high-precision recognition for safe medicine use. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a multi-modal tactile-visual fusion recognition method for Braille bump characters on medicine boxes.

[0006] A multi-modal tactile-visual fusion recognition method for Braille bump characters on medicine boxes according to the present invention includes the following steps:

[0007] Step S1: Real-time collect the three-dimensional topography data and pressure distribution data of Braille bumps through a tactile sensor array;

[0008] Step S2: Synchronously acquire the RGB image, depth information, and surface texture features of the Braille area through a vision sensor, and preprocess the acquired vision data;

[0009] Step S3: Align the tactile data and vision data spatio-temporally, and generate multi-modal joint features through feature-level and decision-level fusion;

[0010] Step S4: Based on the multi-modal joint features, use a machine learning model to complete the semantic parsing and context completion of Braille characters;

[0011] Step S5: Through a multi-channel feedback module, feedback the recognition results through multiple channels including voice, tactile vibration, and a visualization interface.

[0012] Furthermore, the tactile sensor array includes at least one pressure-sensitive element, whose spatial resolution ≥ 0.3 mm, and the force sensitivity range is 0.05 N - 10 N; the pressure-sensitive element includes at least one of a piezoresistive flexible sensor, a capacitive micro-force sensor, or a piezoelectric sensor; the piezoelectric sensor supports the detection of high-frequency vibration signals of 1 kHz - 10 kHz.

[0013] Furthermore, the preprocessing of the vision data includes:

[0014] (1) Perform Braille area segmentation based on an improved U-Net neural network, where the encoder part uses dilated convolution with a dilation rate of 2 to expand the receptive field, and the decoder part introduces a channel attention module (CBAM) to improve the edge segmentation accuracy;

[0015] (2) Use the CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm combined with multi-scale Retinex (Retinal Cortex Theory) for uneven illumination correction;

[0016] (3) Reconstruct the three-dimensional point cloud through the fusion of binocular stereo vision and structured light, and the point cloud density ≥ 1000 points / cm 2 ;

[0017] (4) Use a Gabor filter bank to extract the directional texture features of the surface of Braille bumps.

[0018] Furthermore, the spatio-temporal alignment adopts a microsecond-level synchronization technology based on hardware triggering, specifically including:

[0019] (1) The tactile sensor and the vision sensor achieve microsecond-level synchronization through a hardware trigger signal;

[0020] (2) The improved ICP (Iterative Closest Point) algorithm is used for 3D point cloud registration, and the registration error is optimized to ≤0.1 mm by combining the angle constraint of the point cloud normal vector;

[0021] (3) Motion compensation is performed on the tactile dynamic scanning trajectory and the visual static image through Kalman filtering;

[0022] The improved ICP algorithm specifically includes:

[0023] (a) Calculate the point cloud normal vector based on the covariance matrix of local neighborhood points, with the neighborhood radius ranging from 0.5 mm to 2 mm and the number of neighborhood points ≥15;

[0024] (b) Construct a joint objective function that includes the point cloud position error and the normal vector angle error:

[0025] where p i is the tactile point cloud, q j is the visual point cloud, n i , nj are the corresponding point normal vectors, and λ is the normal vector constraint weight coefficient (0.5 ≤ λ ≤ 2.0);

[0026] (c) Use the singular value decomposition (SVD) method with a damping factor to solve the rotation matrix and translation vector, and the iteration termination condition is that the change rate of the registration error between two iterations <0.01%.

[0027] Furthermore, the feature fusion adopts a hierarchical fusion strategy:

[0028] (1) Feature-level fusion: Dynamically allocate the weights of tactile and visual features through an attention mechanism, and the calculation formula is: where W t , F t are learnable parameters, F fusion is the multi-modal fusion feature vector, F t is the tactile feature vector, F v is the visual feature vector, σ is the Sigmoid function, represents element-wise multiplication, represents feature concatenation;

[0029] (2) Decision-level fusion: Perform confidence fusion on the recognition results of tactile and vision through the D-S evidence theory (Dempster-Shafer evidence theory).

[0030] Furthermore, the machine learning model is a multi-modal neural network architecture, including:

[0031] (1) Tactile Feature Encoder: A temporal feature extraction module that combines 1D-CNN and bidirectional GRU, where the convolutional kernel size of 1D-CNN is 5×1 and the stride is 2;

[0032] (2) Visual Feature Encoder: A multi-scale feature extraction module based on Vision Transformer, including 12 layers of Transformer encoders;

[0033] (3) Cross-modal Interaction Module: Realize the interaction and alignment of tactile and visual features through the cross-attention mechanism;

[0034] (4) Decoder: A sequence decoding module that combines CRF (Conditional Random Field) to output character codes that conform to the GB / T15720-2020 standard.

[0035] Furthermore, the multi-channel feedback module includes:

[0036] (1) Voice Feedback: Support multi-language TTS conversion, and the speech rate can be adaptively adjusted in the range of 120 - 200 words per minute;

[0037] (2) Tactile Vibration Feedback: Generate differentiated vibration patterns, including dynamic combinations with frequencies of 50 - 200Hz and amplitudes of 0.1 - 1.0mm;

[0038] (3) Visualization Interface: Project virtual Braille dot matrices and drug information text superposition through AR glasses, and the three-dimensional height error of the virtual Braille dot matrices is ≤0.05mm.

[0039] Furthermore, a dynamic optimization mechanism is set up, including:

[0040] (1) Sensor Scheduling Strategy Based on Q-Learning: Dynamically switch the dominant sensor according to the environmental light intensity (Lux value) and tactile signal-to-noise ratio (SNR);

[0041] (2) Incremental Learning Module: When the continuous recognition error exceeds 3 times, automatically trigger the online model update process, and the learning rate is set to 0.001;

[0042] (3) Safety Verification Mechanism: Cross-verify the recognition results with the drug database, and start the manual review process when the error exceeds the threshold (1%);

[0043] The online model update process includes:

[0044] (a) Detect unknown Braille variants through an improved One-Class SVM algorithm, and the detection threshold is Mahalanobis distance > 3σ;

[0045] (b) Update the model parameters using an incremental learning strategy, with the sample size for each update ≥ 500, and the self-adaptive adjustment formula for the learning rate being: η t = η0·exp(―γ·t),

[0046] where η0 is the initial learning rate of 0.001, γ is the decay coefficient of 0.005, and t is the number of iterations.

[0047] The multi-modal tactile-visual fusion recognition system for Braille bump characters on medicine boxes includes:

[0048] (1) Tactile perception module: It includes a high-density flexible sensor array, a signal conditioning circuit, and a high-speed ADC conversion unit, where the ADC sampling rate ≥ 200 kHz;

[0049] (2) Visual perception module: It includes a multi-spectral camera (400 - 1000 nm band), a structured light projector, and an image preprocessing unit accelerated by GPU;

[0050] (3) Data processing module: A heterogeneous computing platform equipped with FPGA and NPU, used to run the method, where the NPU computing power ≥ 10 TOPS;

[0051] (4) Feedback execution module: Integrated with a piezoelectric vibrator (response frequency 50 - 200 Hz), a bone conduction speaker, and a Micro-OLED display unit (resolution ≥ 1280×720);

[0052] (5) Communication module: Supports 5G / Bluetooth dual-mode transmission, interacts with the cloud medicine database in real time, and the data transmission encryption complies with AES-256 (Advanced Encryption Standard).

[0053] Furthermore,

[0054] (1) The tactile perception module and the visual perception module can be quickly disassembled and assembled through a magnetic interface, and the disassembly and assembly time ≤ 5 seconds;

[0055] (2) The data processing module is built-in with a security encryption chip, which complies with the HIPAA (Health Insurance Portability and Accountability Act) medical data security standard;

[0056] (3) The feedback execution module supports personalized configuration, and can adjust the feedback intensity according to the user's disability level, including voice volume (0 - 85 dB) and vibration amplitude (0.1 - 1.0 mm).

[0057] Compared with the prior art, the beneficial effects of the present invention are:

[0058] The present invention uses a multi-modal tactile-visual fusion recognition method to effectively solve the problems of low accuracy, poor environmental adaptability, and insufficient multi-modal fusion and collaboration in traditional single-modal recognition technologies. It can accurately collect data such as the three-dimensional topography and pressure distribution of Braille bumps and visual images, and through spatio-temporal alignment and feature fusion, significantly improve the recognition accuracy. At the same time, the dynamic optimization mechanism can adapt to the variants of medicine box Braille, and the multi-channel feedback module can meet the diverse needs of visually impaired people, providing strong support for the safe medication of visually impaired people. Description of the Drawings

[0059] Figure 1 is the main flow chart of multi-modal fusion recognition of the present invention;

[0060] Figure 2 is the flow chart of the improved ICP algorithm of the present invention;

[0061] Figure 3 is the system architecture diagram of multi-modal fusion recognition of the present invention;

[0062] Figure 4 is the interaction diagram of the dynamic optimization mechanism of the present invention; Detailed Description of the Invention

[0063] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0064] As Figures 1 to 4 shown,

[0065] Example 1: Multi-modal Braille Recognition System in the Hospital Pharmacy Scenario

[0066] 1. Scenario Requirements

[0067] The hospital pharmacy needs to quickly and accurately recognize a large number of medicine Braille labels. Especially in an environment where the medicine batches are frequently changed, the labels are worn, or the lighting is complex (such as a mixture of corridor lights and natural light), it is necessary to ensure that the recognition rate is ≥99.9%.

[0068] 2. Technical Implementation Details

[0069] (1) Tactile Sensing Module

[0070] Sensor Array Configuration: A 16×16 piezoresistive flexible sensor array is used, with a spatial resolution of 0.2 mm, covering a 64 mm×64 mm area, which can cover the standard medicine box Braille area (such as the 6×10 dot matrix specified by GB / T15720).

[0071] Dynamic pressure acquisition: The sensor continuously records tactile data at a sampling rate of 100 Hz. When a pressure mutation is detected (such as the bump height difference > 0.15 mm), a secondary scan of the local area is triggered to improve the recognition accuracy of worn Braille.

[0072] (2) Visual perception module

[0073] Hardware configuration:

[0074] Dual-camera system: A main camera (RGB + depth sensor, resolution 12MP) and an auxiliary camera (infrared supplementary light, wavelength 850 nm).

[0075] Structured light projector: Emits periodic stripe light, with a point cloud density reaching 1500 points / cm 2 (exceeding 1000 points / cm 2 ).

[0076] Preprocessing process:

[0077] Illumination correction: The CLAHE algorithm processes the image in blocks, with the block size set to 16×16 pixels, the contrast limit set to 2%, and combined with multi-scale Retinex (scale parameters σ = 15, 30, 60 pixels) to eliminate specular reflection.

[0078] Braille area segmentation: In the improved U-Net, dilated convolution with a dilation rate of 2 is used in the encoder, and the CBAM attention module is introduced in the decoder, with the edge segmentation accuracy improved to 98% (IoU metric).

[0079] (3) Spatiotemporal alignment and fusion

[0080] Hardware synchronization: Triggered by FPGA hardware, the synchronization error between the tactile and visual sensors ≤ 1 μs.

[0081] ICP registration:

[0082] Normal vector calculation: The neighborhood radius is set to 1.2 mm, and the number of neighborhood points ≥ 20 to ensure that the normal vector direction error ≤ 2°.

[0083] Joint objective function: λ is taken as 1.5 to balance the position error and the normal vector constraint, and the registration iteration times ≤ 5 times (the termination condition is that the error change rate < 0.005%).

[0084] (4) Machine learning model

[0085] Cross-modal interaction:

[0086] Tactile encoder: 1D-CNN (convolution kernel 5×1, stride 2) extracts temporal features, and the bidirectional GRU hidden layer is set to 256 units.

[0087] Visual Encoder: Vision Transformer (12 layers, patch size 16×16), outputting multi-scale feature maps.

[0088] Cross-Attention Mechanism: Through QKV self-attention calculation, the number of channels of tactile and visual features is unified to 512 dimensions.

[0089] (5) Feedback and Safety Verification

[0090] Multi-channel Feedback:

[0091] Voice Feedback: Using Chinese-English bilingual TTS, with the speech rate adjusted adaptively (when the ambient noise > 60 dB, the speech rate drops to 150 words per minute).

[0092] Tactile Vibration: Trigger an 80Hz high-frequency vibration (amplitude 0.5mm) when an error is recognized, and a low-frequency 50Hz (0.2mm) when confirmed.

[0093] Safety Verification:

[0094] Compare with the cloud-based drug database (such as the API of the National Medical Products Administration) in real time. If the braille code does not match the database name, initiate the manual review process (the error threshold is set at 1% character error).

[0095] Example 2: Household Portable Braille Recognition Device

[0096] 1. Scenario Requirements

[0097] In a home environment, visually impaired users need to hold the device to quickly recognize the braille on a small medicine box. The device needs to be lightweight (≤300g), low-power (standby ≤ 1W), and support the offline mode (can still work without network).

[0098] 2. Technical Implementation Details

[0099] (1) Tactile Sensing Module

[0100] Micro Sensor Array: Adopt a 4×4 piezoelectric sensor array (single-point force sensitivity 0.01N - 5N), integrated at the front end of the pen-shaped device, and can be manually slid for scanning.

[0101] Dynamic Compensation: Built-in accelerometer to detect the moving speed of the device. When the scanning speed > 5 cm / s, trigger the downsampling mode (tactile sampling rate drops to 50Hz) to avoid motion blur.

[0102] (2) Visual Sensing Module

[0103] Lightweight Configuration:

[0104] Single camera (resolution 5MP, supporting autofocus), paired with a structured light LED (wavelength 940nm).

[0105] Three-dimensional reconstruction adopts monocular vision + structured light, and the point cloud density is optimized to 800 points / cm 2 (Balancing power consumption and accuracy).

[0106] Preprocessing optimization:

[0107] Lighting correction: Set the CLAHE block size to 32×32 pixels to reduce the computational load; Retinex only retains the medium-scale parameter with σ = 15.

[0108] Braille segmentation: Use lightweight U-Net (halve the number of channels), and the inference time ≤ 0.2 seconds.

[0109] (3) Spatiotemporal alignment and fusion

[0110] Low-power synchronization: Trigger through software (the tactile sensor triggers the vision shutter), and the synchronization error ≤ 10 μs.

[0111] ICP simplification algorithm:

[0112] Set the normal vector neighborhood radius to 1.5 mm, and the number of neighborhood points ≥ 10 to reduce the computational complexity.

[0113] The number of iterations is fixed at 3 times, and the registration error is allowed to be ≤ 0.2 mm (balancing accuracy and speed).

[0114] (4) Machine learning model

[0115] Lightweight architecture:

[0116] Tactile encoder: 1D-CNN (convolution kernel 3×1, stride 1) + GRU (128 units).

[0117] Visual encoder: MobileNetV3 (feature extraction layer) + Transformer (6 layers).

[0118] Cross-modal interaction: Use the local attention mechanism (only focus on the local feature area).

[0119] (5) Feedback and adaptive configuration

[0120] Multi-channel feedback:

[0121] Voice feedback: Support dialect recognition (such as Cantonese, dialect TTS), and the speech rate is fixed at 180 words per minute (ambient noise < 40 dB).

[0122] Tactile feedback: Output the encoded pattern through a piezoelectric vibrator (such as "short vibration = correct, long vibration = wrong").

[0123] Personalized settings:

[0124] The user can customize the vibration intensity through tactile buttons (e.g., trigger a 100 Hz vibration with an amplitude of 0.8 mm when braille recognition is incorrect).

[0125] Battery life: continuous operation ≥ 8 hours (tactile + visual hybrid mode).

[0126] (6) Dynamic optimization mechanism

[0127] Incremental learning:

[0128] When detecting unknown braille (e.g., non-standard font), automatically upload it to the cloud and trigger local model update (sample size ≥ 300, learning rate η0 = 0.0005).

[0129] Sensor scheduling:

[0130] In low-light environments (Lux < 50), the system switches to the vision-dominated mode (tactile sampling rate drops to 25 Hz).

[0131] Example 3: Outdoor mobile scenario (e.g., pharmacy counter)

[0132] 1. Scenario requirements

[0133] The outdoor environment has drastic changes in light (e.g., direct sunlight or shadows), and it is necessary to cope with interference such as strong reflections and device shaking. The recognition response time is required to be < 2 seconds.

[0134] 2. Technical implementation details

[0135] (1) Hardware enhancement

[0136] Vision module:

[0137] Install a polarizing filter on the camera to suppress reflections, and increase the power of the structured light projector to 50 mW.

[0138] Adopt a dual-camera (wide-angle + telephoto) and automatically switch the focal length to adapt to different medicine box sizes.

[0139] Tactile module: The sensor array integrates a temperature compensation circuit to avoid force-sensitive drift caused by direct sunlight.

[0140] (2) Algorithm optimization

[0141] Preprocessing enhancement:

[0142] The Retinex algorithm introduces dynamic scale selection: automatically select the σ parameter according to the light intensity (σ = 30 in strong light, σ = 10 in weak light).

[0143] Point cloud reconstruction uses robust random sample consensus (RANSAC) to remove outliers.

[0144] Spatio-temporal alignment:

[0145] Kalman filtering is introduced into motion state estimation: predicting the shaking trajectory of the device through the IMU sensor and compensating for the error of the tactile scanning trajectory.

[0146] (3) Feedback mechanism

[0147] AR visualization: projecting virtual Braille through smart glasses and overlaying information such as the expiration date and usage method of drugs (e.g., highlighting expired drugs in red).

[0148] Voice feedback: supports the environmental noise reduction mode, suppressing background noise through beamforming microphones.

[0149] Key parameters and effect verification

[0150] Hospital scenario:

[0151] Braille recognition accuracy: 99.7% (the recognition rate of worn Braille is increased by 40%).

[0152] System response time: ≤0.8 seconds (including cloud verification).

[0153] Household scenario:

[0154] Low light environment (Lux = 10): recognition rate 98.2% (25% improvement compared to traditional single modality).

[0155] Battery life: continuous operation for 10 hours (tactile + visual hybrid mode).

[0156] Outdoor scenario:

[0157] Under strong reflection (Lux = 10000): point cloud reconstruction error ≤0.15mm, recognition rate 97.5%.

[0158] Through the above specific implementation manners, each module in the technical solution (such as sensor configuration, algorithm parameters, feedback mechanism) has been refined and designed according to different scenario requirements to ensure the practicability and robustness in complex environments.

[0159] The above are only the preferred implementation manners of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-modal tactile-visual fusion recognition method for Braille bump characters on a medicine box, characterized in that, It includes the following steps: Step S1: Real-time collect three-dimensional topography data and pressure distribution data of Braille bumps through a tactile sensor array; Step S2: Synchronously obtain RGB images, depth information and surface texture features of the Braille area through a vision sensor, and preprocess the obtained vision data; Step S3: Align the tactile data and vision data in space and time, and generate multi-modal joint features through feature-level and decision-level fusion; Step S4: Based on the multi-modal joint features, use a machine learning model to complete semantic parsing and context completion of Braille characters; Step S5: Through a multi-channel feedback module, perform multi-channel feedback on the recognition results through voice, tactile vibration and a visualization interface.

2. The multi-modal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, wherein, The tactile sensor array includes pressure-sensitive elements, whose spatial resolution ≥ 0.3 mm and force sensitivity range is 0.05 N - 10 N; The pressure-sensitive elements include at least one of piezoresistive flexible sensors, capacitive micro-force sensors or piezoelectric sensors.

3. The multi-modal tactile-visual fusion recognition method for the braille bump characters of the medicine box according to claim 1, wherein, The preprocessing of the vision data includes: (1) Segment the Braille area based on an improved U-Net neural network, where dilated convolution is introduced in the encoder part to expand the receptive field, and an attention mechanism is added in the decoder part to improve the edge segmentation accuracy; (2) Use the CLAHE algorithm combined with multi-scale Retinex for uneven illumination correction; (3) Reconstruct a three-dimensional point cloud by fusing binocular stereo vision and structured light, with a point cloud density ≥ 1000 points / cm 2 ; (4) Use a Gabor filter bank to extract the microscopic texture features of the surface of Braille bumps.

4. The multi-modal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, characterized in that The space-time alignment adopts a microsecond-level synchronization technology based on hardware triggering, specifically including: (1) The tactile sensor and the vision sensor achieve microsecond-level synchronization through a hardware trigger signal; (2) Use an improved ICP algorithm for three-dimensional point cloud registration, where the registration error is optimized to ≤ 0.1 mm by combining point cloud normal vector constraints; (3) Perform motion compensation on the tactile dynamic scanning trajectory and the vision static image through Kalman filtering; The improved ICP algorithm specifically includes: (a) Calculate the point cloud normal vector based on the covariance matrix of local neighborhood points, the neighborhood radius is 0.5 mm - 2 mm, and the number of neighborhood points ≥ 15; (b) Construct a joint objective function that includes the position error of the point cloud and the angle error of the normal vector: Among them, p i is the tactile point cloud, q j is the visual point cloud, n i , n j are the corresponding point normal vectors, and λ is the normal vector constraint weight coefficient (0.5 ≤ λ ≤ 2.0); (c) Use the singular value decomposition (SVD) method with a damping factor to solve the rotation matrix and translation vector, and the iteration termination condition is that the change rate of the registration error between two iterations < 0.01%.

5. The multi-modal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, wherein, Feature fusion adopts a hierarchical fusion strategy: (1) Feature-level fusion: Dynamically allocate the weights of tactile and visual features through an attention mechanism. The calculation formula is: Among them, W t , F t are learnable parameters, F fusion is a multi-modal fusion feature vector, F t is a tactile feature vector, F v is a visual feature vector, σ is the Sigmoid function, denotes element-wise multiplication, ⊕ denotes feature concatenation; (2) Decision-level fusion: Perform confidence fusion on the recognition results of touch and vision through D-S evidence theory.

6. The multi-modal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, wherein, The machine learning model is a multi-modal neural network architecture, including: (1) Tactile feature encoder: A temporal feature extraction module combining 1D-CNN and bidirectional GRU; (2) Vision feature encoder: A multi-scale feature extraction module based on VisionTransformer; (3) Cross-modal interaction module: Achieve interactive alignment of tactile and vision features through a cross-attention mechanism; (4) Decoder: A sequence decoding module combined with CRF, outputting character codes compliant with the GB / T15720 standard.

7. The multi-modal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, characterized in that, The multi-channel feedback module includes: (1) Voice feedback: Support multi-language TTS conversion and speech rate adaptive adjustment; (2) Tactile vibration feedback: Generate differentiated vibration patterns, including dynamic combinations with frequencies of 50 - 200 Hz and amplitudes of 0.1 - 1.0 mm; (3) Visualization interface: Project virtual Braille dot matrices and drug information text overlays through AR glasses.

8. The multimodal tactile-visual fusion recognition method for the Braille bump characters of the medicine box according to claim 1, characterized in that Set up a dynamic optimization mechanism, including: (1), A sensor scheduling strategy based on reinforcement learning that dynamically switches the dominant sensor according to environmental light intensity and tactile signal-to-noise ratio; (2), An incremental learning module: Automatically trigger the online model update process when an unknown Braille variant is detected; (3), A security verification mechanism: Cross-verify the recognition results with the drug database, and initiate a manual review process when the error exceeds the threshold; The online model update process includes: (a) Detect unknown Braille variants through an improved One-Class SVM algorithm, with a detection threshold of Mahalanobis distance > 3σ; (b) Update the model parameters using an incremental learning strategy. The sample size for each update is ≥500, and the formula for adaptive adjustment of the learning rate is: η t = η0·exp(―γ·t), Among them, the initial learning rate η0 is 0.001, the decay coefficient γ is 0.005, and t is the number of iterations.

9. A multi-modal tactile-visual fusion recognition system for Braille bump characters on a medicine box, characterized in that, Apply the system to implement the method according to any one of claims 1 - 8, and the system includes: (1), A tactile perception module: Comprising a high-density flexible sensor array, a signal conditioning circuit, and a high-speed ADC conversion unit; (2), A visual perception module: Comprising a multispectral camera, a structured light projector, and an image preprocessing unit accelerated by GPU; (3), A data processing module: A heterogeneous computing platform equipped with FPGA and NPU; (4), A feedback execution module: Integrated with a piezoelectric vibrator, a bone conduction speaker, and a Micro-OLED display unit; (5), A communication module: Supports 5G / Bluetooth dual-mode transmission and interacts with the cloud drug database in real time.

10. The multi-modal tactile-visual fusion recognition system for Braille bump characters on a medicine box according to claim 9, wherein (1), The tactile perception module and the visual perception module can be quickly disassembled and assembled through a magnetic interface; (2), The data processing module is built-in with a security encryption chip; (3), The feedback execution module supports personalized configuration and can adjust the feedback intensity according to the user's disability level.

Citation Information

Cited By

  • Portable intelligent device real-time braille point detection method and system for visually impaired people

    CN120823610A

  • Real-time Braille dot detection method and system for portable smart devices for visually impaired individuals

    CN120823610B

  • Visual tactile fusion multi-mode tactile sensor and three-dimensional pressure field reconstruction method

    CN121277350A