Auxiliary navigation method and system based on large language model
Through an auxiliary navigation method based on a large language model, the use of environmental information and line of sight area images to generate audio information, which solves the real-time perception and communication problems of visually impaired people and improves independent living ability.
Patent Information
- Application Number
- CN202411168249.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-08-23
AI Technical Summary
Visually impaired people lack real-time situational understanding and adaptability in their daily lives. Traditional auxiliary tools such as guide dogs have training time and maintenance costs limitations, and an intelligent system is needed to improve independent living capabilities.
Based on the auxiliary navigation method of the large language model, the first device obtains environmental information and line of sight area images, generates text information and converts it into audio information, and uses the second device to perform in-depth analysis and conversion, so as to realize the perception expansion and communication convenience of visually impaired users.
It improves the perception ability and communication convenience of visually impaired users, expands the perception range, realizes instant conversion and playback of environmental information, and transcends the efficiency and convenience of traditional haptic auxiliary devices.
Smart Images

Figure CN119025068B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of deep learning technology, and more specifically, relates to an auxiliary navigation method and system based on a large language model. Background Art
[0002] Visually impaired people face significant challenges in daily life, particularly in activities such as walking, reading, watching TV, and independently completing daily tasks. Traditional assistive devices often lack real-time contextual understanding and adaptability. Guide dogs, while useful, are also limited by training time, cost, and maintenance. Therefore, there is an urgent need for powerful intelligent systems and smart interactive methods to help visually impaired people live better lives. Summary of the Invention
[0003] The purpose of the present disclosure is to provide an assisted navigation method and system based on a large language model to improve the ability of hearing-impaired people to live independently and more easily.
[0004] According to a first aspect of an embodiment of the present disclosure, a large language model-based assisted navigation method is provided, which is applied to a first device and includes:
[0005] determining first text information based on environmental information surrounding the target user and image information within the target user's sight area;
[0006] Sending the first text message and the image information within the target user's sight area to a second device; the second device is a device that has established a connection with the first device; the first text message is used to instruct the second device to generate a second text message based on the first text message and the image information, and convert the second text message into a first audio message;
[0007] The method further comprises responding to the first audio information sent by the second device.
[0008] According to a second aspect of the present disclosure, an audiovisual interaction apparatus is provided, which is applied to a first device and includes:
[0009] A first determining module, configured to determine first text information based on environmental information surrounding the target user and image information within the target user's sight area;
[0010] a first sending module configured to send the first text message and the image information within the target user's visual area to a second device; the second device being a device that has established a connection with the first device; the first text message being configured to instruct the second device to generate a second text message based on the first text message and the image information, and to convert the second text message into a first audio message;
[0011] The first responding module is configured to respond to the first audio information sent by the second device.
[0012] According to a third aspect of the embodiments of the present disclosure, there is provided an auxiliary navigation system based on a large language model, comprising a camera, a microphone, a speaker, a sensor, a data processing module, and a user interface;
[0013] The camera, microphone, speaker, sensor and user interface belong to the first device, and the data processing module and the output module belong to the second device; the data processing module includes a large language model, an intelligent agent, retrieval enhancement generation and a knowledge graph.
[0014] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned auxiliary navigation method based on a large language model are implemented.
[0015] The beneficial effects of the large language model-based assisted navigation method and system provided by the embodiments of the present disclosure are:
[0016] On the one hand, this disclosure can enhance the perception of visually impaired users. By utilizing the user's surroundings and image information within their field of view, the first or second device can convert this non-visual information into text, significantly expanding their perception. Not only can they understand textual information about their surroundings (such as road signs and store names), but they can also gain more detailed information about the environment through image information, such as the location and shape of objects. This significantly enhances their ability to live and travel independently.
[0017] On the other hand, the present disclosure can enhance the convenience of communication. By sending a first text message and an image message to a connected second device (such as a smartphone, smartwatch, or specialized assistive device), and instructing the device to generate a second text message and convert it into the first audio message, visually impaired users can easily access this information. This method is more efficient and convenient than traditional tactile assistive devices (such as Braille displays), and can instantly convert and play information about the surrounding environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1A flowchart of an assisted navigation method based on a large language model provided by an embodiment of the present disclosure;
[0020] Figure 2 An interaction diagram of a large-scale language model assistance system provided by an embodiment of the present disclosure;
[0021] Figure 3 A diagram of the environmental data perception process of a large-scale language model-assisted system provided in one embodiment of the present disclosure;
[0022] Figure 4 A user interaction process of a large-scale language model-assisted system provided by an embodiment of the present disclosure;
[0023] Figure 5 A signaling interaction diagram provided by an embodiment of the present disclosure;
[0024] Figure 6 A component diagram of a large-scale language model assistance system provided by one embodiment of the present disclosure;
[0025] Figure 7 A sequence diagram of a large-scale language model assistance system provided by an embodiment of the present disclosure;
[0026] Figure 8 An activity diagram of a large-scale language model assistance system provided by an embodiment of the present disclosure;
[0027] Figure 9 A use case diagram of various activities of a large-scale language model assistance system provided by an embodiment of the present disclosure;
[0028] Figure 10 A class diagram of a large-scale language model assistance system provided by one embodiment of the present disclosure;
[0029] Figure 11 A state diagram of a large-scale language model assistance system provided by an embodiment of the present disclosure;
[0030] Figure 12 An architectural diagram of a large-scale language model assistance system provided by one embodiment of the present disclosure;
[0031] Figure 13 This is a structural block diagram of an audio-visual interaction device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary detail.
[0033] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, specific embodiments will be described below with reference to the accompanying drawings.
[0034] Please refer to Figure 1 , Figure 1 A flowchart of an assisted navigation method based on a large language model provided in one embodiment of the present disclosure is provided. The method is applied to a first device and includes:
[0035] S101: Determine first text information based on environmental information surrounding a target user and image information within a sight area of the target user.
[0036] In this embodiment, the target user is a user wearing the first device. The target user may be a visually impaired user. Visually impaired users are people who need to use technical means to obtain information about the surrounding environment in order to better navigate and understand the world around them.
[0037] Environmental information refers to information about the target user's current physical environment. This information includes, but is not limited to, location information (such as GPS coordinates), time information (such as date and time), weather conditions, lighting conditions, surrounding buildings, street layout, and pedestrian flow. This information is typically acquired through sensors (such as GPS, ultrasonic sensors, and infrared sensors).
[0038] Image information within the target user's field of view can be acquired via a first device. The first device can be smart glasses with a built-in camera. When the target user wears the smart glasses and goes out, the built-in camera can capture image information of the target user's field of view. This image information may include information such as road signs, scenes, and people, which is crucial for determining the target user's intentions, interests, or needs.
[0039] The first text information refers to preliminary or basic text descriptions or information generated by the smart glasses through preprocessing of the visually impaired user's surroundings and the image information within their field of view. This preprocessing can be done through image feature extraction or sensor data integration. The first text information can be a description of the surrounding environment, such as "There is an obstacle 2 meters ahead" or "The current weather is sunny." The first text information is intended to help visually impaired users or other users requiring assistance better understand and navigate the world around them.
[0040] In this embodiment, various sensors and cameras on the first device (smart glasses) can collect environmental information surrounding a target user and image information within the target user's field of view. The first device can then pre-process this collected environmental and image information to generate first text information, which can help the target user better understand and navigate the world around them.
[0041] S102: Send the first text message and the image information in the target user's field of view to the second device; the second device is a device that has established a connection relationship with the first device; the first text message is used to instruct the second device to generate the second text message based on the first text message and the image information, and convert the second text message into the first audio information.
[0042] In this embodiment, the second device refers to another device that has established a connection relationship with the first device, and can be a server, a computer, a smart phone, etc. The second device can receive the first text information and image information sent by the first device, and perform in-depth analysis and processing on the information to obtain the second text information.
[0043] The first text message not only contains the data to be transmitted, but also some indicative data. This indicative data helps the second device generate a new, more detailed second text message. For example, if the first text message describes "There is a building ahead," and the image information shows that the building is a hospital, the second device can then generate a second text message based on this information, such as "There is a hospital ahead, and there are many pedestrians at the entrance." The first text message provides the basis and direction for the second device to generate the second text message, and the second text message provides a more in-depth or detailed description and summary based on the first text message and image information.
[0044] Because the target user is visually impaired and cannot see images or text in the real world, audio prompts are needed to guide them in completing daily activities. The second device can convert the second text message into a first audio message and send the first audio message to the first device, which then plays it through a built-in speaker or a connected headset.
[0045] S103: Respond to the first audio information sent by the second device.
[0046] In this embodiment, the first device may receive the first audio information sent by the second device via Bluetooth, wireless communication, etc., and play the first audio information to the target user through a built-in speaker or a headset connected thereto.
[0047] From steps S101, S102, and S103, it can be seen that the working principle of the audio-visual interaction method is:
[0048] The first device uses its built-in sensors (such as a microphone, ultrasonic sensor, infrared sensor, gyroscope, etc.) and camera to capture environmental information surrounding the target user, as well as image information within the target user's field of view. The first device uses image processing, natural language processing, and other technologies to analyze the captured sensor data and image information, and combines this with the environmental information to generate a first text message. This first text message typically provides a preliminary description or understanding of the environment within the target user's current field of view, and may include object identification, location description, time information, and other content.
[0049] The first device sends the generated first text message and the image information within the target user's field of view to the second device. The second device is a device that has established a connection relationship with the first device, such as a smartphone, tablet computer or server, which has stronger processing and storage capabilities. After receiving the first text message and image information, the second device uses more complex algorithms and models to comprehensively analyze and process these two types of information. According to the instructions of the first text message, the second device generates a second text message that is more detailed, accurate, and more in line with the needs of the target user than the first text message. The second device converts the second text message into a first audio message and sends it to the first device. The first device can receive the first audio message sent by the second device through Bluetooth, wireless communication, etc., and play the first audio message to the target user through the built-in speaker or the headphones connected thereto.
[0050] As can be seen from the above, the present disclosure can, on the one hand, enhance the perception of visually impaired users. By utilizing the environmental information surrounding the target user and the image information within their field of view, the first or second device can convert this non-visual information into textual information, thereby significantly expanding the visually impaired user's perception range. Not only can they understand the textual information in their current environment (such as road signs and store names), but they can also understand more environmental details such as the location and shape of objects through image information, which greatly enhances their ability to live and travel independently.
[0051] On the other hand, the present disclosure can enhance the convenience of communication. By sending a first text message and an image message to a connected second device (such as a smartphone, smartwatch, or specialized assistive device), and instructing the device to generate a second text message and convert it into the first audio message, visually impaired users can easily access this information. This method is more efficient and convenient than traditional tactile assistive devices (such as Braille displays), and can instantly convert and play information about the surrounding environment.
[0052] In one embodiment of the present disclosure, an assisted navigation method based on a large language model further includes:
[0053] In response to receiving the first request information from the target user, sending the first request information to the second device;
[0054] The first request information is used to instruct the second device to obtain the third text information according to the first request information and convert the third text information into the second audio information;
[0055] The second device responds to the second audio information sent by the second device.
[0056] In this embodiment, the first request information refers to a request information sent by a target user (e.g., a visually impaired user) through some means (e.g., voice command, gesture recognition, physical button, etc.). This information can instruct the second device to perform a specific task or operation, such as querying the weather, sending a message, or navigating to a specific location.
[0057] The third text information is the text information obtained by the second device after receiving the first request information and processing. The information may include specific content that meets the user's request, such as weather forecast, navigation route description, message text, etc.
[0058] The second audio information is audio information obtained by converting the third text information by the second device. Similar to the first audio information, the second audio information can also be played to the target user through headphones or speakers so that the visually impaired user can understand and receive it.
[0059] In this embodiment, the first device first receives a first request message from a visually impaired user, which specifically indicates the user's needs or query content. Then, the first device forwards this request message to the second device with which a connection has been established. Based on the received request message, the second device autonomously performs corresponding data retrieval, processing or calculation tasks to obtain third text information that matches the user's request. The second device also uses text-to-speech conversion technology to convert the third text information into a second audio message that is easy for the visually impaired user to understand. Finally, the visually impaired user obtains the required information by listening to the second audio message played by the first device, and can make corresponding responses or further operations based on the content, thereby achieving an efficient and convenient audio-visual interactive experience.
[0060] From the above, it can be concluded that the present disclosure significantly improves the efficiency and convenience of interaction between visually impaired users and devices. Through the automated text acquisition and audio conversion process, it reduces the complexity of user operations, ensures the instant transmission and accurate reception of information, and promotes smoother and barrier-free information exchange.
[0061] In one embodiment of the present disclosure, determining the first text information based on the environmental information surrounding the target user and the image information within the target user's sight area includes:
[0062] Analyze the environmental information around the target user to obtain the first type of text;
[0063] Perform text extraction on the image information within the target user's sight area to obtain the second type of text;
[0064] The first text information is determined according to the first type of characters and the second type of characters.
[0065] In this embodiment, the first type of text is obtained by the first device parsing the environmental information around the target user. The environmental information can be collected by an ultrasonic sensor, an infrared sensor, or other sensors. For example, an ultrasonic sensor can measure the distance and shape of surrounding objects by emitting ultrasonic waves and receiving their echoes, which is extremely important for the first device to identify obstacles in the environment, determine spatial layout, and other information. Although the ultrasonic sensor itself does not directly generate text information, its measurement data can be used by the first device to infer environmental characteristics and then generate descriptive text descriptions as part of the first type of text. The infrared sensor can detect infrared radiation emitted by objects. Through the infrared sensor, the first device can sense whether there are other organisms around the target user and their activity status, thereby generating corresponding text descriptions.
[0066] The second type of text is obtained by extracting text from images within the target user's field of view. This image information specifically refers to the area where the user's current line of sight is focused, and can be captured by the camera of a head-mounted device (such as smart glasses). Text extraction technologies (such as optical character recognition (OCR)) are used to identify text within images, such as text in books, instructions on product packaging, and information on distant road signs. This second type of text more accurately reflects the specific content of the user's current attention.
[0067] After obtaining the first and second types of text, the first device can determine the first text information according to certain rules or algorithms. This process may include steps such as text screening, deduplication, and sorting to ensure that the generated text information is comprehensive and accurate, and can effectively reflect the current environment and visual focus of the target user. From the above, it can be concluded that this embodiment achieves a comprehensive analysis of the target user's surrounding environment and visual focus by integrating multiple sensor data and image recognition technology. This not only improves the accuracy of environmental perception, but also enhances the ability to capture user focus. The first text information generated by this embodiment is comprehensive and accurate, and can effectively help users understand the current environment and enhance interactive experience and convenience.
[0068] In one embodiment of the present disclosure, sending the first text information and the image information within the sight area of the target user to the second device includes:
[0069] The first text information and the image information in the target user's sight area are converted into information in a preset format, and the information in the preset format is sent to the second device.
[0070] In this embodiment, the first device needs to convert the first text message and the image information within the target user's field of view into information in a predetermined format. This predetermined format can be a data format that the second device can directly recognize and process, such as JSON, XML, or a data packet of a specific protocol. This embodiment ensures the integrity and consistency of the information through format conversion, reducing the risk of data loss or corruption during transmission.
[0071] After the format conversion is completed, the first device sends the generated preset format information to the second device. This step may involve the selection and implementation of a network communication protocol, such as TCP / IP, HTTP, or other wireless communication protocols to ensure that the information can be reliably transmitted to the second device.
[0072] As can be seen from the above, this embodiment ensures the integrity and consistency of information through pre-set format conversion, effectively reducing the risk of loss and corruption during data transmission. Furthermore, the selection of an appropriate network communication protocol enhances the reliability of information transmission, allowing the first device to efficiently and securely transmit user environment information and image information to the second device, improving the efficiency of multi-device collaboration and the user experience.
[0073] In one embodiment of the present disclosure, generating second text information according to first text information and image information includes:
[0074] The first text information and the image information are input into a large language model (LLM) to obtain a correlation analysis result between the first text information and the image information and the second text information.
[0075] In this embodiment, the trained large-scale language model can process the first text information and the image information in the target user's field of view to obtain the second text information. The large-scale language model can be OpenAI's GPT series model, Meta's important open source large-scale language model, etc.
[0076] In this embodiment, a large-scale language model is used to process both the first textual information and the image information. This innovative approach significantly enriches the dimensionality and depth of information processing. By simultaneously feeding two different types of information (textual information and image information) into the large-scale language model, the model, drawing on its vast knowledge base and complex algorithmic logic, analyzes the potential connections and correlations between the two types of information. Specifically, the large-scale language model first performs semantic understanding and parsing of the first textual information, while simultaneously using image recognition technology to identify and understand the content of the image information. The large-scale language model then integrates these two types of information through a deep cross-analysis, generating a second textual information that captures the essence of the original textual information while incorporating the relevance of the image information. This process not only enhances the richness and diversity of information but also generates new insights and perspectives.
[0077] From the above, it can be concluded that this embodiment integrates text and image information through a large language model, which not only improves the depth and breadth of information processing, but also can generate second text information that better meets user needs.
[0078] In one embodiment of the present disclosure, an assisted navigation method based on a large language model further includes:
[0079] In response to a similarity between the multiple images within the sight line of the target user within the first time period being greater than a first threshold, any one of the multiple images within the sight line of the target user is used as image information within the sight line of the target user.
[0080] In this embodiment, the first duration refers to a specific time interval, during which the first device can continuously capture image information within the target user's field of view. The first threshold is a preset value used to determine whether multiple images within the target user's field of view are similar. When the similarity is greater than the first threshold, it indicates that the image changes little, and any one of them can be used as representative image information. For example, when the target user needs to read a book, the book can be placed within the field of view that can be captured by the smart glasses. When the smart glasses determine that the image information within the target user's field of view within the first duration is highly similar and basically unchanged, it indicates that the user may want to read a book. At this time, the smart glasses can send the captured image information of the book to the second device, and the second device converts the text in the image into audio information and sends it to the first device, which plays it to the target user, thereby realizing the reading function for visually impaired users.
[0081] In one embodiment of the present disclosure, the first request information includes a first trigger instruction and / or third audio information from a target user to the first device.
[0082] In this embodiment, the first request information is a request signal or message sent by the target user during the interaction between the target user and the first device. The first request information is used to initiate a series of operations, such as sending relevant data or requesting a service to the second device. The first trigger instruction is an action instruction issued by the target user to the first device in some way, such as pressing a button on the first device, a gesture instruction, etc. The third audio information refers to the sound data input into the first device by the target user through voice or other means. The third audio information may include the target user's verbal instructions, questions or other voice content, which is used to instruct the first device how to process information or transmit specific instructions to the second device.
[0083] As can be seen from the above, this embodiment enables the first device to accurately capture the user's intent and initiate the corresponding operation by receiving the target user's first request information, including the first trigger instruction and / or the third audio information. This not only enhances the user experience, but also improves the efficiency and accuracy of interaction with the second device, ensuring that the target user can quickly and accurately obtain the required service or information.
[0084] In one embodiment of the present disclosure, reference Figure 2 , Figure 2 An interaction diagram of a large-scale language model assistance system provided by one embodiment of the present disclosure.
[0085] The cameras and sensors built into the smart glasses first capture environmental and image data, then send them to the data processing unit for data preprocessing. This process includes sensing the environment, understanding semantics, and generating contextual information (i.e., generating the first text message). The control system built into the smart glasses then sends this preprocessed data to the large model (short for large language model). The large model uses Retrieval-Augmented Generation (RAG) and knowledge graphs to conduct in-depth analysis of this data to obtain the second text message. The large model then converts the second text message into voice information, and the voice output module (i.e., the speakers built into the smart glasses) provides voice feedback or voice guidance.
[0086] After receiving the voice feedback, the user can submit interaction requirements to the smart glasses through the user interface. The user interface sends the voice instructions or the user's trigger instructions to the data processing unit for preprocessing. Then the control system built into the smart glasses sends the preprocessed data to the large model for in-depth analysis and output voice guidance. After receiving the voice guidance, the built-in speaker of the smart glasses plays the voice to the user, and the user can complete daily activities.
[0087] In one embodiment of the present disclosure, reference Figure 3 , Figure 3The environmental data perception process for large-scale language model-assisted systems.
[0088] First, smart glasses can capture environmental data through built-in sensors, and then send the environmental data, text data and sensor data to the data processing unit for processing. The data processing unit inputs the pre-processed data into a large language model for data analysis. After the data analysis is completed, context information is obtained. The large language model can convert the context information into audio output information and provide audio feedback to the user through the smart glasses.
[0089] In one embodiment of the present disclosure, reference Figure 4 , Figure 4 User interaction process for large language model-assisted systems.
[0090] First, the user sends an interaction request to the user interface, and the user interface transmits the data request to the data processing unit. The data processing unit sends the user instruction to the large language model for processing to obtain context feedback information. At the same time, the large language model processes the context feedback information to obtain response data, and uses the audio output unit to convert the response data into audio feedback information. The audio feedback information provides audio feedback to the user through the smart glasses.
[0091] In one embodiment of the present disclosure, reference Figure 5 , the signaling interaction between the first device (smart glasses) and the second device (the server containing the large language model) is as follows:
[0092] A1: Acquire environmental information and image information;
[0093] Smart glasses can use built-in sensors to obtain environmental information around the target user and use built-in cameras to obtain image information within the target user's field of view.
[0094] A2: Preprocessing the environmental information and image information to obtain first text information;
[0095] The data processing unit pre-processes the environmental information and the image information to perceive the environment, understand the semantics, and generate context information to obtain the first text information.
[0096] A3: Send the first text message and image message;
[0097] The data processing unit sends the pre-processed first text information and image information to the large language model of the server, and the large language model further analyzes and processes the first text information and image information to provide the user with more accurate voice guidance.
[0098] A4: Further processing the first text information and the image information to obtain second text information;
[0099] The large language model further processes the first text information and image information to obtain response data. At the same time, RAG enhances the large language model by retrieving the latest relevant information, and combines the retrieved data with the generated response to obtain the second text information.
[0100] A5: converting the second text information into the first audio information;
[0101] The server converts the second text information into first audio information.
[0102] A6: Send the first audio message;
[0103] The server sends the first audio information to the smart glasses.
[0104] A7: broadcast the first audio message;
[0105] The speaker in the smart glasses can broadcast the first audio information sent by the server.
[0106] A8: Receive the first request information;
[0107] After the target user learns about the surrounding environment through the first audio information, he or she may perform relevant actions on the smart glasses, such as sending a voice message, pressing a button on the smart glasses, etc. The smart glasses may receive the first request information from the target user for the smart glasses.
[0108] A9: Send the first request information;
[0109] The smart glasses send the first request information (such as a voice command) to the server.
[0110] A10: Obtain third text information according to the first request information;
[0111] The server calls the information of the relevant software according to the voice instruction to obtain the third text information. For example, the server can call the map software to find the best walking path for the visually impaired user.
[0112] A11: converting the third text information into second audio information;
[0113] The server converts the third text information into second audio information.
[0114] A12: Send the second audio message;
[0115] The server sends the second audio information to the smart glasses.
[0116] A13: broadcast the second audio message;
[0117] The smart glasses broadcast the second audio information received from the server, so that the target user can perform activities such as walking, reading, and running.
[0118] In one embodiment of the present disclosure, the audio-visual interaction method is applied to a large language model-assisted system. The system comprises: smart glasses, a software system (large language model, intelligent agent, search enhancement generation, knowledge graph), etc. The smart glasses are equipped with a camera, a microphone, a speaker, and a distance measurement sensor (e.g., a lidar sensor, an ultrasonic sensor), etc. The component diagram of the system is shown in FIG. Figure 6 .
[0119] exist Figure 6 In the smart glasses, the built-in camera is used to capture images, which are then processed by the visual processing unit and sent to the decision engine for analysis. The decision engine can control the camera to capture images in real time. Smart glasses can also use built-in sensors to detect obstacles. If the obstacle detection module identifies an obstacle, it will send the obstacle data to the decision engine, which will convert the obstacle information into audio data and play it to the user through the speaker. The decision engine can also receive external commands and send the output audio to the speaker through the audio output module, which will play the audio. Smart glasses can also use a microphone to receive audio data and send the audio data to the decision engine for analysis and processing through the audio processing unit. Intelligent agents are software systems that enable interaction between users, large language models, and hardware devices.
[0120] Environmental data provided by devices such as sensors can enrich the content of knowledge graphs, which can then accurately identify objects based on the environmental data provided by sensors and combined with user preference data. Retrieval-augmented generation is a method for optimizing the output of large language models, enabling them to reference authoritative knowledge bases outside of the training data source before generating a response. Natural language processing (NLP) engines can be used to perform text analysis and generate contextual information using large language models. NLP engines can also perform speech synthesis based on contextual information, which can be played back to the user through the smart glasses' speakers.
[0121] In one embodiment of the present disclosure, a sequence diagram of a visually impaired user interacting with a large language model assistance system is provided. Figure 7 .
[0122] First, the system is initialized, and then the intelligent agent component is activated, which in turn activates sensors, cameras, microphones, and speakers. Visually impaired users can input, such as navigation, through a user interface (including controls for user interaction with the system, such as requesting specific information or help through voice commands or a tactile interface on the glasses). After the smart glasses capture the user input (e.g., voice signals), the intelligent agent analyzes the input information and sends the information to a large language model (large language model). The large language model uses RAG to retrieve relevant information, obtains data from external data sources (e.g., maps), combines this data with the response, and generates a response for output. The intelligent agent provides output (e.g., audio guidance) based on the response, so that the visually impaired user can hear the relevant audio information.
[0123] For example, the input provided by the visually impaired user through the user interface is to read text, the smart glasses send the processed input to the smart agent, the smart agent instructs the camera to capture an image, the camera sends the collected image data to the smart agent, the smart agent sends the image data to the large language model, the large language model performs optical character recognition (OCR), and converts the recognized text into speech and sends it to the smart agent, the smart agent provides output (such as audio reading) based on the speech, and the visually impaired user can then hear the audio information of the relevant book. In one embodiment of the present disclosure, the activity diagram of the visually impaired user performing various activities is referenced. Figure 8 .
[0124] Figure 8 In the process, the system needs to be initialized first, then the sensors, cameras, microphones, and speakers are activated, and then the user can provide input. When the input is for navigation, the system processes the output, detects obstacles, and provides real-time guidance. When the input is for reading text, the system processes the input, captures images, performs OCR, converts text to language, and provides audio reading. When the input is for watching TV, the system processes the input, synchronizes with the TV program, and provides audio description. When the input is for recognizing the road environment, the system processes the input, analyzes visual data, recognizes road signs and signals, and provides road safety guidance. When the input is for running or participating in a marathon, the system processes the input, detects obstacles, plans the path, and provides real-time feedback.
[0125] In one embodiment of the present disclosure, see Figure 9 , Figure 9 Use case diagram for each activity.
[0126] exist Figure 9In the system, a visually impaired user (who can be blind) first initializes the system, activating sensors, a large language model, and an intelligent agent. When the visually impaired user's activity involves running or participating in a marathon, the intelligent agent is required for real-time feedback and route planning. When the visually impaired user's activity involves identifying the road environment, a large language model is used to analyze visual data and recognize road signs and signals. External data sources are imported through RAG to generate response data. The intelligent agent can also provide road safety guidance. When the visually impaired user's activity involves navigation and obstacle avoidance, sensors can be used to detect obstacles, while the intelligent agent provides guidance. When the visually impaired user's activity involves watching TV, a large language model can be used to provide audio descriptions of the visual content. When the visually impaired user's activity involves reading text, image data can be captured through a camera, processed by optical character recognition (OCR), and then converted to speech using a large language model. Finally, the speech data is played back to the visually impaired user.
[0127] In one embodiment of the present disclosure, reference Figure 10 , Figure 10 This is a class diagram for a large-scale language model-assisted system. IntelligentGlasses represents smart glasses. The camera class is represented by captureImage():Image; the microphone class is represented by captureAudio():Audio; the speaker class is represented by outputAudio(audio: String):void; and the sensor class is represented by detectObstacle():Boolean. IntelligentAgent represents an intelligent agent, including the classes manageInteraction(input: String):String and makeDecision(data:String):String.
[0128] RetrievalAugmentedGeneration represents retrieval augmented generation, which includes the classes retrievalnformation(query: String): String and generateResponse(data: String):String.
[0129] LargeLanguageModel represents a large language model, which contains classes such as processinput(input:String): String and generateResponse(data: String): String.
[0130] TextAnalysis represents text analysis and includes classes such as performOCR(image: lmage): String.
[0131] SpeechSynthesis represents speech synthesis and includes the class convertTextToSpeech(text:String): Audio.
[0132] KnowledgeGraph represents a knowledge graph and includes classes such as getContextualData(query:String): String.
[0133] In one embodiment of the present disclosure, reference Figure 11 , Figure 11 This is a state diagram for a large language model-assisted system. It first waits for commands, and then, when the system receives a command from the user, it begins processing. The processing includes processing navigation, road recognition, and providing guidance; processing road recognition, environmental analysis, and providing guidance; processing operation, detecting obstacles, planning paths, and providing guidance; processing reading, performing OCR, and converting text to speech; and processing TV program viewing and depicting visual content. For detailed processing procedures, refer to Figure 9 Activity use case diagram.
[0134] In one embodiment of the present disclosure, reference Figure 12 , Figure 12 This is the architecture diagram of the large-scale language model auxiliary system.
[0135] The hardware components include smart glasses, which include sensors, speakers, cameras, and microphones.
[0136] Software components include intelligent agents, knowledge graphs, retrieval-enhanced generation, and large language models.
[0137] External systems include tracking systems, external data sources, etc.
[0138] The intelligent agent is a software component that manages interactions between users, large language models, and hardware. It implements decision-making algorithms for dynamic task execution and interacts with other components to provide more comprehensive data. Large language models are AI models based on deep learning technology and can be trained on diverse datasets, including text, image, and audio data, to ensure robustness and accuracy in various scenarios. Large language models also have natural language understanding and generation capabilities, enabling real-time interaction and emotional awareness assistance for users.
[0139] Retrieval-augmented generation (RAG) can enhance the output of large language models by retrieving the latest information relevant to the input content. The large language model then combines the retrieved data with the generated response to provide more accurate assistance to users. Knowledge graphs are a semantic network-based knowledge representation and management technology that can structure knowledge related to visually impaired activities, including environmental data, object recognition, and user preferences.
[0140] External systems can help software components make better decisions.
[0141] In one embodiment of the present disclosure, an audio-visual interaction method can assist visually impaired users with walking assistance. For example, a visually impaired user named Xiao Ming wears smart glasses (a first device) and goes out. The smart glasses acquire surrounding environmental information, such as the current location on a street and good weather, as well as image information of the road ahead and pedestrians within their field of view. The smart glasses pre-process this information to generate a first text message, such as "You are on a street. The road ahead is flat and there are a few pedestrians." The first text message and image information are then sent to a server (a second device) containing a large language model. The server further processes this information to generate a more detailed second text message, such as "You are on XX street. The road ahead is flat 10 meters ahead and there are three pedestrians walking slowly." This text message is converted into a first audio message and sent back to the smart glasses. The smart glasses broadcast the first audio message, and Xiao Ming, upon hearing it, wishes to know if there is a bus stop ahead. He sends a voice request to the smart glasses. The smart glasses then send the request to the server. Based on the request, the server retrieves information from relevant mapping software to generate a third text message, such as "There is a bus stop 50 meters ahead and the route is XXX." This text message is then converted into a second audio message and sent to the smart glasses. The smart glasses broadcast the second audio information, and Xiao Ming can walk according to the voice prompts of the smart glasses.
[0142] In one embodiment of the present disclosure, an audio-visual interaction method can help visually impaired users to read.
[0143] For example, a visually impaired user named Xiaohong wants to read a new novel. She puts on smart glasses (the first device) and places the book within her field of vision. The smart glasses capture the book's image and the current environment, such as whether the room is moderately lit and quiet. The smart glasses pre-process this information to generate a first text message, such as "The room is quiet, and there is an image of a book page within the field of vision." The first text message and the book page image are then sent to a server (the second device) that contains a large language model. The server further processes this information using the large language model, identifying the text on the book page and generating detailed second text message, such as "This page reads: The protagonist begins his adventure in a mysterious town..." This second text message is converted into a first audio message and sent back to the smart glasses. The smart glasses then play the first audio message. After hearing a section of the first audio message, Xiaohong wants to know the meaning of a certain word. She sends a voice request to the smart glasses, which then sends the request to the server. Based on the request, the server retrieves the relevant meaning as a third text message, such as "The word 'mysterious' here means elusive or unknown." The second device then converts the third text message into a second audio message and sends it to the smart glasses. The smart glasses broadcast a second audio message to help Xiaohong better understand the reading content and continue to enjoy the fun of reading.
[0144] In one embodiment of the present disclosure, a user may utilize a first device and a second device to perform daily activities such as cooking, shopping, and using home appliances.
[0145] Corresponding to the above embodiment, an assisted navigation method based on a large language model is described. Figure 13 This is a structural block diagram of an audio-visual interaction device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 13 The audio-visual interaction device 20 is applied to a first device and includes: a first determination module 21, a first sending module 22 and a first response module 23.
[0146] The first determining module 21 is configured to determine the first text information based on the environment information surrounding the target user and the image information within the target user's sight area;
[0147] A first sending module 22 is configured to send the first text message and image information within the target user's visual area to a second device; the second device is a device that has established a connection with the first device; the first text message is used to instruct the second device to generate a second text message based on the first text message and image information, and convert the second text message into a first audio message;
[0148] The first responding module 23 is configured to respond to the first audio information sent by the second device.
[0149] In one embodiment of the present disclosure, an audio-visual interaction device 20 further includes: a second response module;
[0150] The second response module is configured to send the first request information to the second device in response to receiving the first request information of the target user;
[0151] The first request information is used to instruct the second device to obtain the third text information according to the first request information and convert the third text information into the second audio information;
[0152] The second device responds to the second audio information sent by the second device.
[0153] In one embodiment of the present disclosure, the first sending module 22 is specifically configured to:
[0154] The first text information and the image information in the target user's sight area are converted into information in a preset format, and the information in the preset format is sent to the second device.
[0155] In one embodiment of the present disclosure, the first sending module 22 is specifically configured to:
[0156] Inputting the first text information and the image information into a large language model;
[0157] The first text information is also used to instruct the large language model to combine the generated response with the retrieval enhancement generation and the content of the knowledge graph to obtain the second text information.
[0158] In one embodiment of the present disclosure, an audio-visual interaction device 20 further includes a third response module;
[0159] The third response module is configured to, in response to the similarity of the multiple images within the target user's sight area within the first time period being greater than a first threshold, use any one of the multiple images within the target user's sight area as image information within the target user's sight area.
[0160] In one embodiment of the present disclosure, the first request information includes a first trigger instruction and / or third audio information from a target user to the first device.
[0161] The present disclosure also provides an assisted navigation system based on a large language model, comprising a camera, a microphone, a speaker, a sensor, a data processing module, and a user interface;
[0162] The camera, microphone, speaker, sensor and user interface belong to the first device, and the data processing module belongs to the second device; the data processing module includes a large language model, intelligent agent, retrieval enhancement generation and knowledge graph.
[0163] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.
[0164] The computer-readable storage medium can be the internal storage unit of the smart device in any of the aforementioned embodiments, such as the smart device's hard drive or memory. The computer-readable storage medium can also be an external storage device of the smart device, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both the internal storage unit of the smart device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the smart device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.
[0165] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.
[0166] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the smart devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed intelligent devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or it can be an electrical, mechanical or other form of connection.
[0168] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present disclosure.
[0169] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0170] The above are only specific embodiments of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or replacements within the technical scope disclosed in this disclosure, and such modifications or replacements should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. An assisted navigation method based on a large language model, characterized in that: Applied to a first device, comprising: determining a first text message based on environmental information surrounding the target user and image information within the target user's field of view; the first text message including not only data to be transmitted but also indicative data, the indicative data being used to help the second device generate a second text message; Sending the first text message and the image information within the target user's sight area to a second device; the second device is a device that has established a connection with the first device; the first text message is used to instruct the second device to generate a second text message based on the first text message and the image information, and convert the second text message into a first audio message; responding to the first audio information sent by the second device; Determining the first text information based on the environment information surrounding the target user and the image information within the target user's sight area includes: Analyze the environmental information around the target user to obtain the first type of text; Perform text extraction on the image information within the target user's sight area to obtain the second type of text; determining first text information according to the first type of characters and the second type of characters; Generating the second text information according to the first text information and the image information includes: inputting the first text information and the image information into a large language model; The first text information is also used to instruct the large language model to combine the generated response with the retrieval enhancement generation and the content of the knowledge graph to obtain the second text information.
2. The assisted navigation method based on a large language model according to claim 1, characterized in that: Also includes: In response to receiving the first request information of the target user, sending the first request information to the second device; The first request information is used to instruct the second device to obtain third text information according to the first request information and convert the third text information into second audio information; Respond to the second audio information sent by the second device.
3. The assisted navigation method based on a large language model according to claim 1, characterized in that: Sending the first text information and the image information within the target user's sight area to the second device includes: The first text information and the image information in the target user's sight area are converted into information in a preset format, and the information in the preset format is sent to the second device.
4. The assisted navigation method based on a large language model according to claim 1, characterized in that: Also includes: In response to a similarity between the multiple images within the sight line of the target user within the first time period being greater than a first threshold, any one of the multiple images within the sight line of the target user is used as image information within the sight line of the target user.
5. The assisted navigation method based on a large language model according to claim 2, characterized in that: The first request information includes a first trigger instruction and / or third audio information from a target user to the first device.
6. An audio-visual interactive device, characterized in that: Applied to a first device, comprising: a first determining module configured to determine first text information based on environmental information surrounding a target user and image information within a visual area of the target user; the first text information including not only data to be transmitted but also indicative data, the indicative data being used to assist the second device in generating a second text information; The first determining module is specifically configured to: Analyze the environmental information around the target user to obtain the first type of text; Perform text extraction on the image information within the target user's sight area to obtain the second type of text; determining first text information according to the first type of characters and the second type of characters; a first sending module configured to send the first text message and the image information within the target user's visual area to a second device; the second device being a device that has established a connection with the first device; the first text message being configured to instruct the second device to generate a second text message based on the first text message and the image information, and to convert the second text message into a first audio message; Generating the second text information according to the first text information and the image information includes: inputting the first text information and the image information into a large language model; The first text information is further used to instruct the large language model to combine the generated response with the search enhancement generation and the content of the knowledge graph to obtain the second text information; The first responding module is configured to respond to the first audio information sent by the second device.
7. An auxiliary navigation system based on a large language model, characterized in that: Includes cameras, microphones, speakers, sensors, data processing modules and user interfaces; The camera, the microphone, the speaker, the sensor and the user interface belong to the first device, and the data processing module belongs to the second device; the data processing module includes a large language model, an intelligent agent, retrieval enhancement generation and a knowledge graph.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Scene character interactive understanding system for visually impaired people
CN114168104A
Blind person assisting method and system based on artificial intelligence, terminal and storage medium
CN117982321A
Image description method and device, equipment, storage medium and product
CN118467776A