Online target identification method based on lightweight deep learning model

Through customized lightweight deep learning models and multi-scene data acquisition and processing, the efficient identification problem of visually impaired devices in complex environments is solved, the convenience and safety of visually impaired people are improved, and social interaction is enhanced.

CN120298764APending Publication Date: 2025-07-11林佳涛
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351231.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing blindness equipment is difficult to efficiently identify complex and diverse obstacles and road conditions information on portable devices, and the lightweight model cannot meet the high-precision identification needs of visually impaired people in multiple scenarios, and lacks integrated optimization for specific scenarios.

Method used

Using a customized lightweight deep learning model, combined with deep separation convolution and pruning technology, we build visual impairment auxiliary equipment suitable for visual impairment, collect data in real time and preprocess it through a high-definition camera, and use low-power Bluetooth module to transmit image data to the main control chip for feature extraction and classification. We combine GPS module and visual impairment auxiliary APP to achieve real-time object recognition and personalized services in multiple scenarios.

Benefits of technology

It achieves rapid and accurate goal recognition under limited hardware resources, improves the convenience, autonomy and safety of life for visually impaired people, and enhances interaction and social support between family members and volunteers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005326071580000011
    Figure HDA0005326071580000011
Patent Text Reader

Abstract

The invention discloses an online target identification method based on a lightweight deep learning model, and belongs to the field of artificial intelligence and visual impairment assistance, and the method comprises a customized lightweight model construction step: designing a lightweight deep learning model suitable for visual impairment assistance equipment through employing a depth separable convolution and pruning technology, the number of model parameters and calculation complexity are reduced, pre-training is carried out on a large-scale general image data set, and then fine tuning is carried out based on visual impairment-related travel, shopping, home and social scene data; the method is adaptive to visual impairment auxiliary equipment such as a vision field navigation cap, target identification in travel, shopping, home and social scenes can be quickly and accurately realized under the condition of limited hardware resources, the actual requirements of visual impairment people are comprehensively met, and the convenience, autonomy and safety of life of the visual impairment people are improved; and an innovative scene fusion and intelligent interaction mechanism provides more intelligent and personalized auxiliary services for visually impaired people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and visually impaired assistance technology, and particularly to an online target recognition method based on a lightweight deep learning model. Background Art

[0002] The visually impaired population faces many challenges in daily travel, shopping, home life, social interaction and other scenarios. In terms of travel, existing guiding devices are difficult to accurately identify complex and diverse obstacles and road conditions, and cannot provide comprehensive and reliable navigation guidance for the visually impaired. When shopping, it is difficult for the visually impaired to independently obtain information such as the types of goods, prices, and the areas where the goods are located. In home life, there are difficulties in identifying the production dates of goods and medicines, and daily necessities. In social scenarios, there is also a lack of effective means to assist in identifying acquaintances and the emotions of others. Current target recognition technologies require a large amount of computing resources for complex models and cannot run smoothly on devices such as the portable "Vision Navigation Hat"; while lightweight models are difficult to meet the requirements of the visually impaired for high-precision recognition in multiple scenarios, and lack integration and optimization for these specific scenarios, and cannot provide one-stop effective assistance for the visually impaired population. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the present invention aims to develop an online target recognition method based on a lightweight deep learning model, which is deeply adapted to devices such as the "Vision Navigation Hat", and can quickly and accurately achieve target recognition in travel, shopping, home life, and social scenarios under limited hardware resource conditions, fully meeting the actual needs of the visually impaired, and improving the convenience, autonomy and safety of their lives.

[0004] Technical Solution: To solve the above technical problems, according to one aspect of the present invention, more specifically, an online target recognition method based on a lightweight deep learning model, which is applied to a visually impaired assistance device, includes the following steps:

[0005] S1. Customized lightweight model construction step: Use depthwise separable convolution and pruning techniques to design a lightweight deep learning model suitable for a visually impaired assistance device, reduce the number of model parameters and computational complexity, pre-train on a large-scale general image dataset, and then fine-tune based on data related to travel, shopping, home life, and social scenarios for the visually impaired;

[0006] S2. Multi-scenario data collection and processing step: Use the high-definition camera equipped with the visually impaired assistance device to collect image data in real time in complex environments and different scenarios such as crowded streets, large shopping malls, indoor homes, and social gathering places, perform preprocessing such as normalization, contrast enhancement, and noise removal on the collected data, and use data augmentation techniques such as random rotation, flipping, translation, and adding noise;

[0007] S3. Full-scenario online recognition steps: The camera captures the surrounding environment images in real time, transmits them to the host main control chip via the low-power Bluetooth module. The main control chip calls the lightweight deep learning model, and uses feature extraction and classification algorithms to perform real-time analysis and accurate recognition on obstacles, traffic signs, and the location of blind paths in the travel scenario, the types of goods, brands, price tags, and the location of the goods area in the shopping scenario, the production dates, shelf lives, remaining quantities and locations of goods and medicines in the home scenario, and the facial features of acquaintances and the emotional states of others in the social scenario. Then, the recognition results are converted into voice information through the voice module and conveyed to the user through the directional sound transmission earphone;

[0008] S4. Scenario fusion and intelligent interaction steps: Integrate the GPS module positioning information, and with the help of the visually impaired assistance APP, build a real-time data interaction system among the visually impaired assistance device, visually impaired users, family members, and volunteers, provide location-based personalized target recognition services for visually impaired users, and at the same time facilitate family members to view travel trajectories, locations, and recognition scenario information, as well as remote assistance from volunteers.

[0009] Furthermore, the visually impaired assistance device is the "Vision Navigation Cap", and the "Vision Navigation Cap" is equipped with a camera, a Bluetooth module, a main control chip, a voice module, a directional sound transmission earphone, a GPS module, a directional sound transmission earphone, and a storage module for storing data.

[0010] Furthermore, during the target recognition and processing process, the lightweight deep learning model dynamically adjusts the recognition strategy and parameters according to different scenario characteristics and requirements, and automatically records image data for subsequent optimization for targets that are difficult to accurately recognize.

[0011] Furthermore, the visually impaired assistance APP supports family members and volunteers to view the location information, travel trajectories, and target recognition results of visually impaired users in real time, and can send voice or text prompts to the "Vision Navigation Cap", and also supports information sharing and communication among multiple users.

[0012] Furthermore, during device initialization, each component of the "Vision Navigation Cap" automatically performs self-checks, and establishes a connection with the user's mobile device via Bluetooth, and starts the visually impaired assistance APP to complete the pairing and data synchronization of the device and the application program.

[0013] Furthermore, during the data collection and transmission process, the camera collects image data at a preset frequency of 10 frames per second, and after format conversion and compression processing, it is transmitted to the main control chip via Bluetooth. When the Bluetooth connection is abnormal, it automatically reconnects and gives a voice prompt to the user.

[0014] Furthermore, during the result feedback and interaction process, the "Vision Navigation Cap" supports the user to interact with the device through voice commands, and the device makes corresponding answers and navigation guides according to the positioning information and recognition results.

[0015] Furthermore, during the data storage and model optimization process, the stored image data, recognition results, and interaction records are regularly analyzed and sorted, representative sample data is extracted to retrain and optimize the lightweight deep learning model, and the model structure and parameters are adjusted according to user feedback and scenario changes.

[0016] The beneficial effects of an online target recognition method based on a lightweight deep learning model according to the present invention are as follows:

[0017] (1) The present invention is deeply adapted to visually impaired assistance devices such as the "Vision Navigation Cap", and can quickly and accurately achieve target recognition in travel, shopping, home, and social scenarios under limited hardware resources, fully meeting the actual needs of visually impaired people and improving the convenience, autonomy, and safety of their lives.

[0018] (2) The innovative scenario integration and intelligent interaction mechanism not only provides more intelligent and personalized assistance services for visually impaired people, but also strengthens the connection and interaction between them and their family members and volunteers, building a solid bridge for the visually impaired group to integrate into society and having significant social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described in detail below with reference to the drawings and specific implementation methods.

[0020] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0021] The present invention will be described in detail below with reference to the drawings and embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0022] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0023] (1) Device initialization: When the "Vision Navigation Cap" is turned on, each component (camera, Bluetooth module, GPS module, main control chip, voice module, directional sound transmission earphone, etc.) automatically performs an initialization self-check to ensure that the device is in a normal working state. At the same time, the "Vision Navigation Cap" establishes a connection with the user's mobile device (such as a mobile phone) through Bluetooth, starts the visually impaired assistance APP, and completes the pairing and data synchronization of the device and the application program.

[0024] (2) Data Acquisition and Transmission: During the user's daily activities, the camera continuously and real-time collects image data of the surrounding environment at a preset frequency (such as 10 frames per second). After the collected image data undergoes preliminary format conversion and compression processing, it is sent to the main control chip of the host through the Bluetooth module at a stable and high transmission rate. During this process, if the Bluetooth connection encounters an abnormality, the device will automatically attempt to reconnect and inform the user of the current connection status through voice prompts.

[0025] (3) Target Recognition and Processing: After receiving the image data, the main control chip immediately loads the trained lightweight deep learning model. The model first extracts features from the image, using multiple convolutional neural networks and pooling layers to extract the key feature information of the targets in the image. Then, through the fully connected layer and classifier, the extracted features are analyzed and classified to determine whether there are specific targets in scenarios such as traveling, shopping, at home, and socializing in the image, and to determine the specific categories and attributes of the targets. During the recognition process, the model will dynamically adjust the recognition strategy and parameters according to the characteristics and requirements of different scenarios. For example, in the traveling scenario, it focuses on the recognition of targets such as obstacles and road signs; in the shopping scenario, it gives priority to recognizing features such as the appearance and labels of goods. For some targets that are difficult to accurately recognize, the model will automatically record the relevant image data and conduct key learning and improvement during the subsequent model optimization process.

[0026] (4) Result Feedback and Interaction: Once the main control chip completes the target recognition, the recognition result will be immediately transmitted to the voice module. The voice module converts the recognition result into clear and concise voice information according to the preset voice templates and rules. For example, in the traveling scenario, the voice prompt may be "There is a step 2 meters ahead, please walk carefully"; in the shopping scenario, the voice prompt may be "You are currently in the food area. There is milk 1 meter ahead, and the price is 10 yuan". The converted voice information is accurately conveyed to the user through the directional sound transmission earphone to ensure that the user can clearly hear the relevant prompts. At the same time, the "Vision Navigation Cap" also supports the user to interact with the device through voice commands. For example, the user can ask "Is there a pharmacy nearby" through voice, and the device will provide the corresponding answer and navigation guidance according to the current positioning information and recognition result, referring to Figure 1 .

[0027] (V) Data Storage and Model Optimization: During the operation of the device, all the collected image data, recognition results, and user interaction records are stored in real-time in the storage module of the host (or synchronously stored in the cloud server through the network). These data will serve as an important basis for subsequent model optimization and improvement. Analyze and organize the stored data regularly (such as weekly or monthly), extract representative sample data from it, and use it for retraining and optimizing the lightweight deep learning model. At the same time, adjust and optimize the structure and parameters of the model according to the user's feedback and changes in the actual usage scenario, continuously improving the recognition accuracy and adaptability of the model.

[0028] (VI) Multi-User Collaboration and Remote Assistance For the family members and volunteers of visually impaired users, they can view the location information, travel trajectory, and current target recognition results of the user through the visually impaired assistance APP in real-time. When family members or volunteers find that the user may encounter difficulties or dangers, they can send voice or text prompts to the "Vision Navigation Cap" through the APP to provide remote assistance and guidance to the user. In addition, the visually impaired assistance APP also supports information sharing and communication among multiple users, forming a mutually supportive community environment. For example, visually impaired users can share their travel experiences and problems encountered on the APP, and other users or volunteers can provide corresponding suggestions and help, further improving the quality of life and social support network of the visually impaired group.

[0029] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. An online object recognition method based on a lightweight deep learning model, which is applied to a visually impaired assistance device, and is characterized in that, It includes the following steps: S1. Customized lightweight model construction step: Use depthwise separable convolution and pruning techniques to design a lightweight deep learning model suitable for visually impaired assistance devices, reduce the number of model parameters and computational complexity, pre-train on a large-scale general image dataset, and then fine-tune based on data of travel, shopping, home, and social scenarios related to the visually impaired; S2. Multi-scenario data collection and processing step: Use the high-definition camera equipped on the visually impaired assistance device to collect image data in real time in complex environments and different scenarios such as crowded streets, large shopping malls, indoor homes, and social gathering places, perform preprocessing such as normalization, contrast enhancement, and noise removal on the collected data, and adopt data augmentation techniques such as random rotation, flipping, translation, and adding noise; S3. Full-scenario online recognition step: The camera captures the surrounding environment image in real time, transmits it to the host main control chip through the low-power Bluetooth module, the main control chip calls the lightweight deep learning model, and uses feature extraction and classification algorithms to perform real-time analysis and accurate recognition on obstacles, traffic signs, and blind path positions in the travel scenario, product types, brands, price tags, and product area positions in the shopping scenario, production dates, shelf lives, remaining quantities and positions of daily necessities, and facial features of acquaintances and emotional states of others in the social scenario, and converts the recognition results into voice information through the voice module and conveys it to the user through the directional sound transmission earphone; S4. Scenario fusion and intelligent interaction step: Integrate the positioning information of the GPS module, and with the help of the visually impaired assistance APP, build a real-time data interaction system among the visually impaired assistance device, visually impaired users, family members, and volunteers, provide location-based personalized target recognition services for visually impaired users, and at the same time facilitate family members to view travel trajectories, locations, and recognition scenario information, as well as remote assistance from volunteers.

2. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: The visually impaired assistance device is the "Vision Navigation Cap", and the "Vision Navigation Cap" is equipped with a camera, a Bluetooth module, a main control chip, a voice module, a directional sound transmission earphone, a GPS module, a directional sound transmission earphone, and a storage module for storing data.

3. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: During the target recognition processing, the lightweight deep learning model dynamically adjusts the recognition strategy and parameters according to different scenario characteristics and requirements, and automatically records the image data of the targets that are difficult to accurately recognize for subsequent optimization.

4. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: The visually impaired assistance APP supports family members and volunteers to view the location information, travel trajectories, and target recognition results of visually impaired users in real time, and can send voice or text prompts to the "Vision Navigation Cap", and also supports information sharing and communication among multiple users.

5. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: During device initialization, each component of the "Vision Navigation Cap" automatically performs self-checking, and establishes a connection with the user's mobile device through Bluetooth, and starts the visually impaired assistance APP to complete the pairing and data synchronization of the device and the application program.

6. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: During the data collection and transmission process, the camera collects image data at a preset frequency of 10 frames per second, and transmits it to the main control chip through Bluetooth after format conversion and compression processing. When the Bluetooth connection is abnormal, it automatically reconnects and gives a voice prompt to the user.

7. The online target recognition method based on a lightweight deep learning model according to claim 1, wherein: During the result feedback and interaction process, the "Horizon Pilot Cap" supports users to interact with the device through voice commands, and the device makes corresponding responses and navigation guidance based on the positioning information and recognition results.

8. The online target recognition method based on a lightweight deep learning model according to claim 1, characterized in that: During the data storage and model optimization process, the stored image data, recognition results, and interaction records are regularly analyzed and sorted out, representative sample data is extracted to retrain and optimize the lightweight deep learning model, and the model structure and parameters are adjusted according to user feedback and scenario changes.