Complex environment recognition and intelligent interaction system for blind navigation
By combining the blind navigation system with object detection algorithm and large language model, the existing system's shortcomings in intelligence and interactivity are solved, and more accurate and humanized blind travel assistance is achieved, meeting the navigation and help seeking needs in the complex environment of the blind group.
Patent Information
- Application Number
- CN202510046596.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing blind travel assistance systems have shortcomings in terms of intelligence and interactivity, and it is difficult to meet the travel needs of blind people, especially in navigation and help seeking in complex environments.
It adopts a complex environment recognition and intelligent interaction system for blind navigation, combined with object detection algorithms and large language models, through the collaborative work of client devices and background servers, it recognizes objects and marking points in the environment in real time, provides voice broadcasting and interaction functions, and improves navigation accuracy and humanization level.
It realizes more accurate road conditions and blind path information monitoring, provides more humanized voice broadcasting and interaction functions, improves the convenience and safety of blind people's travel, and meets the higher interactive requirements and more humanized travel assistance needs of blind people.
Smart Images

Figure CN120101786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of navigation for the blind, and in particular to a complex environment recognition and intelligent interaction system for navigation for the blind. Background Art
[0002] Pedestrian navigation assistance equipment is an important facility in the field of navigation. It can be widely used in traffic guidance and daily travel. By combining with traffic facilities along the route and walking conditions, pedestrian navigation assistance equipment also has broad application prospects in pedestrian travel planning, saving pedestrian travel time, and improving walking comfort. At present, pedestrian navigation assistance equipment has poor compatibility with special groups. For the travel assistance needs of a large number of blind people, some companies have proposed some blind travel assistance solutions based on wearable devices and mobile terminals, but there are still some problems that have not been effectively solved. This method is aimed at the travel needs of the blind group, explores complex environment recognition and intelligent interaction methods for blind navigation, and attempts to solve the problems of difficulty in travel, poor direction perception, and inconvenience in seeking help for the blind group.
[0003] Existing travel assistance systems for the blind are mainly divided into travel assistance devices and applications based on mobile smart terminals: travel assistance devices are often based on wearable hardware facilities, and have problems such as heavy weight, complex carrying, and failure to fully consider the user's own conditions when designing. They are generally expensive and have a limited audience. Mobile smart terminal-based systems are some applications installed on mobile smart terminals that are carried daily, which solve the problems of convenience and usage cost, but have defects and shortcomings such as low intelligence, insufficient interactivity, and difficulty in meeting the travel needs of the blind.
[0004] At present, in the field of travel assistance methods for people with disabilities, especially the blind, mainstream scholars mostly use image recognition, artificial intelligence and other fields to study feasible methods. However, there are few studies related to intelligent interaction and humanized guidance for blind travel assistance, and the domestic and foreign achievements have certain shortcomings. For example, navigation broadcasts are limited to the "several meters" that ordinary people are accustomed to instead of "several steps", and the status of existing blind path facilities is not fully explored. Even though the methods and models proposed by some scholars have a certain competitiveness in the effect of blind travel assistance, the demand for higher interaction requirements and more humanized travel assistance is still difficult to meet. Summary of the invention
[0005] The purpose of the present invention is to provide a complex environment recognition and intelligent interaction system for blind navigation, which is a blind travel assistance navigation implementation strategy that can recognize complex environments and provide intelligent interaction. It gives full play to the advantages of front-end and back-end collaboration and integrates multiple algorithms to achieve more accurate road conditions and blind path information monitoring and more humanized voice broadcasting and interactive functions, providing better travel assistance services for the visually impaired group.
[0006] To achieve the above purpose, the technical solution provided by the present invention is: a complex environment recognition and intelligent interaction system for blind navigation, including a client device and a backend server, wherein the client device is equipped with a speaker, a microphone, a pedometer, an acceleration sensor and a camera at a fixed position, and has a built-in target detection algorithm and a map navigation module with marked points; the backend server has a built-in large visual model and a large language model, and has a computing function;
[0007] The specific implementation of the system includes the following steps:
[0008] S1. The client device obtains image data of the surrounding environment in real time through the camera, including continuous multi-frame images, and uses the target detection algorithm to process the image data to obtain the object category and size information labels in the image data; the client device uses the map navigation module with the marked points to identify the location information labels of the marked points within the surrounding area of its positioning in real time;
[0009] S2, transmitting the object category and size information label and the mark point location information label in a specified data format to the large language model of the backend server to generate corresponding instructions, and after the content of the instructions is transmitted back to the client device, the object category, size and location information are output in the form of voice, and the user can choose to communicate with the client device to achieve interaction;
[0010] S3, calling the pedometer and acceleration sensor to identify a stable step cycle, collecting the distance data of the user walking in the stable step cycle, and calculating the average step length of the user;
[0011] S4. The client device measures the distance between the object and the user through monocular vision technology, and combines the average step length to output the number of steps required to reach the object through voice, where the number of steps is obtained by dividing the distance by the average step length.
[0012] Further, in step S1, the target detection algorithm includes using the YOLOv8 target detection algorithm with the CBAM attention mechanism added and calling the visual big model API to assist in accurate image recognition;
[0013] The network structure of the YOLOv8 target detection algorithm includes three parts: Backbone, Neck and Head; the Backbone part is a feature extraction network that extracts useful information from the input image, and adopts a fast-implemented CSP bottleneck module (i.e., C2f module), and improves the feature extraction capability through the bottleneck block Bottleneck Block and the fast spatial pyramid pooling layer (i.e., SPPF module); the Neck part is located between the Backbone and the Head, and adopts a path aggregation network-feature aggregation network structure (i.e., PAN-FAN structure), and performs multi-scale feature fusion through the path aggregation network (i.e., PAN) and the feature aggregation network (i.e., FAN); the Head part is responsible for predicting the category of the target, and adopts a decoupled head structure without an anchor frame to separate the classification and detection heads;
[0014] The YOLOv8 target detection algorithm adds a CBAM attention mechanism to the Neck part. For the input image features of different scales extracted by its CSP LayerDarknet backbone network, the CBAM attention mechanism is used to perform attention calculations on the features in the channel and spatial dimensions, distinguish the features between different channels and capture the spatial structure in the image, and pass the image features of different scales processed by the CBAM attention mechanism to the Neck part for feature fusion;
[0015] When the confidence level of the result obtained by the target detection algorithm is lower than the set threshold, the visual big model API is called to transmit the image data acquired by the camera in real time to the visual big model for recognition, and the object category and size information label is given.
[0016] Further, in step S1, when a mark point appears within the positioning range of the client device, the map navigation module of the client device will generate a location information tag of the mark point.
[0017] Further, in step S2, the large language model summarizes the information tags and outputs the text content in the form of voice to form an interaction with the user. The background server is mounted on the NVIDIANeMo framework. The specified format specifies the field information when the client device and the background server transmit data. The original binary data is Base64 encoded, and then the data is serialized into JSON format. The hash value is calculated to ensure integrity, and then the whole containing the data and the hash value is compressed, including the "image_data, image_info, audio_data, landmark_name" fields, where image_data, image_info, audio_data, and landmark_name represent image data, image information, voice information, and marker information, respectively.
[0018] Further, in step S3, an autocorrelation analysis algorithm is used to reduce the influence of conditions that may affect acceleration changes during the travel process on the calculation of the step value;
[0019] First, the three-axis acceleration α of the client device x 、a y , α z Take the modulus and calculate the overall acceleration α, and judge the user's motion state based on the standard deviation σ of the overall acceleration:
[0020]
[0021]
[0022] Where u is the overall acceleration sequence of the user's step cycle {α 1 ,α 2 ,...,α k ,...,α N}, α k represents the kth overall acceleration, N represents the number of overall accelerations, and the user state is determined to be idle or walking according to the value of the standard deviation of the overall acceleration. When the standard deviation of the overall acceleration is less than the set threshold, it can be judged that the user's activity amplitude at this moment is small, and the movement state at this moment is determined to be idle, otherwise it is walking;
[0023] Since the user walks continuously and rhythmically, the overall acceleration has obvious periodic changes over time. This periodic change can be used to calculate the autocorrelation between the overall acceleration of the current step cycle and the previous step cycle, and the user's motion state can be further judged based on the size of the autocorrelation. When the user walks, the acceleration sensor will continuously record data. After calculating the overall acceleration, the autocorrelation coefficient of the overall acceleration x(m,t) is calculated using the following formula, where m is the step cycle and t is the sampling period:
[0024]
[0025] Where u(m,t) and σ(m,t) represent the mean and standard deviation of the acceleration sequence {a(p), a(p+1), ..., a(p+1-t}) in the current step cycle m within the sampling period t; u(m+t,t) and σ(m+t,t) are the values of u(m,t) and σ(m,t) after a sampling period t. p is an auxiliary time parameter, which can help calculate the appropriate cycle length from 0 to avoid calculation omissions. a(p) is the overall acceleration at time p. a(m+p) is the overall acceleration at time m+p, and a(m+p+t) is the overall acceleration at time m+p+t. When the sampling period t is close to the user's step period m, the walking states represented by a(m+p) and a(m+p+t) in the formula can be considered consistent, so the value of x(m,t) is close to 1. However, the step frequency of different users or the same user at different times is inconsistent, so t is a variable, and the framing algorithm needs to be used to dynamically determine t, that is, the value interval of t is selected as [t min , t max ], t min Represents the minimum value of the sampling period t, t max Represents the maximum value of the sampling period t, and calculates the new autocorrelation coefficient ρ(m,t):
[0026]
[0027] The t value when ρ(m, t) reaches its maximum value is the user's stable step cycle. In this cycle, the ratio of the walking distance to the pedometer step count is the average step length.
[0028] Further, in step S4, the monocular vision technology uses a triangular transformation method to measure the distance between the object and the user. There are two cases:
[0029] The first one is when the object is directly in front of the camera:
[0030] Through inverse trigonometric functions, we can get the angle between the camera optical axis and the Y axis of the world coordinate system, and the angle between the line connecting the measured point and the origin of the image coordinate system and the camera optical axis. By subtracting the two angles, we can get the angle ∠Υ between the line connecting the measured point and the origin of the image coordinate system and the Y axis of the world coordinate system. By dividing the camera height H by the tangent value of ∠Υ, we can get the required distance in this case.
[0031] The second case is when the object being measured is not directly in front of the camera:
[0032] The point to be measured is recorded as Q, and its projection point in the image coordinate system is recorded as Q'. The distance between Q' and the camera is obtained from the first case, thereby obtaining the geometric distance from the origin of the camera coordinate system to the projection of the point to be measured on the Y axis of the world coordinate system. Since the triangle formed by the origin of the camera coordinate system, the point to be measured and the projection of the point to be measured on the Y axis of the world coordinate system has a similar relationship with the projection of the triangle in the image coordinate system, the required distance in this case can be calculated by measuring the distance from the origin in the camera coordinate system to the projection of the point to be measured, combined with the above similarity relationship.
[0033] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0034] 1. The present invention combines the target detection algorithm with the large language model, which brings multiple advantages to the system and realizes the integration of text and visual information. The target object identified by the target detection algorithm is associated with the natural language text generated by the large language model, providing richer and more accurate environmental perception and understanding. On this basis, the intelligent semantic understanding and generation capabilities of the large language model enable the system to generate detailed descriptions, instructions and suggestions based on the identified target objects, improving the communication effect and the accuracy of information transmission. The self-evolution of deep learning enables the system to continuously optimize the algorithm through interaction and feedback with the user, providing smarter and more personalized auxiliary functions, providing users with richer, more accurate and smarter auxiliary functions, improving user experience and usage effects, and bringing substantial improvements to the field of blind travel.
[0035] 2. The present invention has a wide range of applications, not only limited to indoor navigation, but also includes outdoor navigation, travel exploration, and daily life assistance. In response to the needs of blind users, the system of the present invention provides a full range of navigation and environmental perception services, enabling them to better integrate into society and live independently: In terms of outdoor navigation, the system of the present invention uses target detection algorithms and large language models to identify and describe various target objects in the outdoor environment. Through real-time target recognition and language generation, the system can provide detailed navigation route guidance and also provide a description of the surrounding environment.
[0036] 3. In order to solve the problem of perception difficulty when using meters as the unit of length in existing navigation systems, the present invention combines the pedometer, acceleration sensor functions and monocular vision technology to change the broadcast unit from "meters" to "steps", which is more in line with the usage habits of blind users and improves the humanization level of the system, thereby providing blind users with a more professional navigation experience that better meets their needs, while enhancing user safety protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of deploying the system and collecting data in a navigation device for the blind.
[0038] Figure 2 It is a geometric diagram of the distance measurement method when the object being measured is directly in front of the camera.
[0039] Figure 3 It is a geometric diagram of the distance measurement method when the object being measured is not directly in front of the camera.
[0040] Figure 4 It is a schematic diagram for realizing personalized reporting of step units. DETAILED DESCRIPTION
[0041] The present invention is further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0042] like Figure 1 As shown, this embodiment discloses a complex environment recognition and intelligent interaction system for blind navigation, including a client device and a backend server, wherein the client device is equipped with a speaker, a microphone, a pedometer, an acceleration sensor, and a camera at a fixed position, and has a built-in target detection algorithm and a map navigation module with marked points; the backend server has a built-in visual large model and a large language model, and has a computing function;
[0043] The specific implementation of the system includes the following steps:
[0044] S1. The client device obtains image data of the surrounding environment in real time through the camera, including continuous multi-frame images, and uses the target detection algorithm to process the image data to obtain the object category and size information labels in the image data; the client device uses the map navigation module with the marked points to identify the location information labels of the marked points within the surrounding area of its positioning in real time; the details are as follows:
[0045] The target detection algorithm includes using the YOLOv8 target detection algorithm with the CBAM attention mechanism added and calling the visual large model API to assist in accurate image recognition;
[0046] The network structure of the YOLOv8 target detection algorithm includes three parts: Backbone, Neck and Head; the Backbone part is a feature extraction network that extracts useful information from the input image, and adopts a fast-implemented CSP bottleneck module (i.e., C2f module), and improves the feature extraction capability through the bottleneck block Bottleneck Block and the fast spatial pyramid pooling layer (i.e., SPPF module); the Neck part is located between the Backbone and the Head, and adopts a path aggregation network-feature aggregation network structure (i.e., PAN-FAN structure), and performs multi-scale feature fusion through the path aggregation network (i.e., PAN) and the feature aggregation network (i.e., FAN); the Head part is responsible for predicting the category of the target, and adopts a decoupled head structure without an anchor frame to separate the classification and detection heads;
[0047] The YOLOv8 target detection algorithm adds a CBAM attention mechanism to the Neck part. For the input image features of different scales extracted by its CSP LayerDarknet backbone network, the CBAM attention mechanism is used to perform attention calculations on the features in the channel and spatial dimensions, distinguish the features between different channels and capture the spatial structure in the image, and pass the image features of different scales processed by the CBAM attention mechanism to the Neck part for feature fusion;
[0048] When the confidence level of the result obtained by the target detection algorithm is lower than the set threshold, the visual big model API is called to transmit the image data acquired by the camera in real time to the visual big model for recognition, and the object category and size information label is given.
[0049] When a marked point appears within the positioning range of the client device, the map navigation module of the client device will generate a location information tag of the marked point.
[0050] S2. The object category and size information label and the mark point location information label are transmitted to the large language model of the backend server in a specified data format to generate corresponding instructions. After the content of the instructions is transmitted back to the client device, the object category, size and location information are output in the form of voice. The user can choose to communicate with the client device to achieve interaction, as follows:
[0051] The large language model aggregates the information tags and outputs the text content in the form of voice to form an interaction with the user. The background server is mounted on the NVIDIANeMo framework. The specified format specifies the field information when the client and the background server transmit data. The original binary data is Base64-encoded, and then the data is serialized into JSON format. The hash value is calculated to ensure integrity, and then the whole containing the data and hash value is compressed. The transmitted information includes "image_data, image_info, audio_data, landmark_name", which respectively represent "image data, image information, voice information, and landmark information". In this example, image_data represents the detected image data, which is a Base64-encoded string; image_info represents the basic information of the image, including width, height, and image format; audio_data represents the collected voice information, which is a Base64-encoded string; landmark_name represents the landmark information recognized by the map navigation module.
[0052] One of the transmission examples is shown in Table 1:
[0053] Table 1: Examples of some transmission information
[0054]
[0055] In the second row of Table 1, the height of the transmitted image information is 1280 pixels, the width is 1616 pixels, the partial string encoded by Base64 is / 9j / 4ShJRXhpZgAATU0AKgAAAAg…, the partial string encoded by Base64 of the transmitted voice information is "UklGRhQAAABXQV…, because there is no marker point in the surrounding area of the client device positioning, there is no marker point information, and landmark_name is displayed as "No"; in the third row of Table 1, the height of the transmitted image information is 474 pixels, the width is 632 pixels, the partial string encoded by Base64 is / 9j / 4AAQSkZJRgABAQEAAAAAAAD…, the partial string encoded by Base64 of the transmitted voice information is "UklGRiAAAABXQVZF…, there is a marker point "Garbage Station" in the surrounding area of the client device positioning, and landmark_name is displayed as "Garbage Station".
[0056] S3, calling the pedometer and acceleration sensor to identify a stable step cycle, collecting the distance data of the user walking in the stable step cycle, and calculating the average step length of the user, as follows:
[0057] The autocorrelation analysis algorithm is used to reduce the influence of the conditions that may affect the acceleration change during the travel process on the calculation of the step value;
[0058] First, the three-axis acceleration α of the client device x 、a y , α z Take the modulus and calculate the overall acceleration α, and judge the user's motion state based on the standard deviation σ of the overall acceleration:
[0059]
[0060]
[0061] Where u is the overall acceleration sequence of the user's step cycle {α 1 ,α 2 ,...,α k ,...,α N}, α k represents the kth overall acceleration, N represents the number of overall accelerations, and the user state is determined to be idle or walking according to the value of the standard deviation of the overall acceleration. When the standard deviation of the overall acceleration is less than the set threshold, it can be judged that the user's activity amplitude is small at this moment, so the motion state at this moment is determined to be idle, otherwise it is walking. In this embodiment, it is considered that the probability of the pedestrian state being idle during the period when the standard deviation of the overall acceleration is less than 0.5 is relatively high, so the motion state at this moment is determined to be idle to reduce the possibility of erroneous counting, otherwise it is walking.
[0062] Since the user walks continuously and rhythmically, the overall acceleration has obvious periodic changes over time. This periodic change can be used to calculate the autocorrelation between the overall acceleration of the current step cycle and the previous step cycle, and the user's motion state can be further judged based on the size of the autocorrelation. When the user walks, the acceleration sensor will continuously record data. After calculating the overall acceleration, the autocorrelation coefficient of the overall acceleration x(m,t) is calculated using the following formula, where m is the step cycle and t is the sampling period:
[0063]
[0064] Where u(m,t) and σ(m,t) represent the mean and standard deviation of the acceleration sequence {a(p), a(p+1), ..., a(p+1-t}) in the current step cycle m within the sampling period t; u(m+t,t) and σ(m+t,t) are the values of u(m,t) and σ(m,t) after a sampling period t. p is an auxiliary time parameter, which can help calculate the appropriate cycle length from 0 to avoid calculation omissions. a(p) is the overall acceleration at time p. a(m+p) is the overall acceleration at time m+p, and a(m+p+t) is the overall acceleration at time m+p+t. When the sampling period t is close to the user's step period m, the walking states represented by a(m+p) and a(m+p+t) in the formula can be considered consistent, so the value of x(m,t) is close to 1. However, the step frequency of different users or the same user at different times is inconsistent, so t is a variable, and the framing algorithm needs to be used to dynamically determine t, that is, the value interval of t is selected as [t min , t max ], t min Represents the minimum value of the sampling period t, t max Represents the maximum value of the sampling period t, and calculates the new autocorrelation coefficient ρ(m,t):
[0065]
[0066] The t value when ρ(m, t) reaches its maximum value is the user's stable step cycle. In this cycle, the ratio of the walking distance to the pedometer step count is the average step length.
[0067] S4. The client device measures the distance between the object and the user through monocular vision technology, and outputs the number of steps required to reach the object in voice, based on the average step length. The number of steps is obtained by dividing the distance by the average step length.
[0068] The monocular vision technology uses a triangular transformation method to measure the distance between an object and a user. There are two situations:
[0069] The first one is when the object is directly in front of the camera:
[0070] like Figure 2 As shown, let the camera height be H, O c The distance to Q' is f, and the component of f on the Y axis of the world coordinate system is f y According to the tangent relationship, the calculation formulas for the angle ∠α between the camera optical axis and the Y axis of the world coordinate system and the angle ∠β between the line connecting the measured point and the origin of the image coordinate system and the camera optical axis can be obtained:
[0071]
[0072] Among them, O represents the origin of the image coordinate system, Oc Represents the origin of the camera coordinate system, O w is the origin of the world coordinate system, M is the intersection of the camera optical axis and the Y axis of the world coordinate system, P is the measurement point on the ground plane, O w P is the required vertical distance.
[0073] The P coordinate is (u 0 ,v 0 ), the coordinates of the projection point P' in the image coordinate system are (u,v), then OP'=vv 0 , from which we can get the calculation formula of the vertical distance of point P:
[0074]
[0075] Among them, ∠Υ is obtained by subtracting ∠α and ∠β.
[0076] The second case is when the object being measured is not directly in front of the camera:
[0077] like Figure 3 As shown, the vertical distance O of the known target point w P and ∠γ, from which we can get O c P:
[0078]
[0079]
[0080] Where Q represents the distance measurement point in the world coordinate system, and Q' is its projection point in the image coordinate system.
[0081]
[0082] P′Q′=(uu 0 )×dy
[0083] Therefore, the horizontal distance PQ of the target point can be solved by the following formula:
[0084]
[0085] like Figure 4 As shown, the distance between the blind person's position and the target point can be measured.
[0086] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A complex environment recognition and intelligent interaction system for blind navigation, characterized in that: The system comprises a client device and a backend server, wherein the client device is equipped with a speaker, a microphone, a pedometer, an acceleration sensor and a camera at a fixed position, and has a built-in target detection algorithm and a map navigation module with markers; The backend server has a built-in large visual model and a large language model, and has computing capabilities; The specific implementation of the system includes the following steps: S1. The client device obtains image data of the surrounding environment in real time through the camera, including continuous multi-frame images, and uses the target detection algorithm to process the image data to obtain the object category and size information labels in the image data; the client device uses the map navigation module with the marked points to identify the location information labels of the marked points within the surrounding area of its positioning in real time; S2, transmitting the object category and size information label and the mark point location information label in a specified data format to the large language model of the backend server to generate corresponding instructions, and after the content of the instructions is transmitted back to the client device, the object category, size and location information are output in the form of voice, and the user can choose to communicate with the client device to achieve interaction; S3, calling the pedometer and acceleration sensor to identify a stable step cycle, collecting the distance data of the user walking in the stable step cycle, and calculating the average step length of the user; S4. The client device measures the distance between the object and the user through monocular vision technology, and combines the average step length to output the number of steps required to reach the object through voice, where the number of steps is obtained by dividing the distance by the average step length.
2. The complex environment recognition and intelligent interaction system for blind navigation according to claim 1, characterized in that: In step S1, the target detection algorithm includes using the YOLOv8 target detection algorithm with the CBAM attention mechanism added and calling the visual big model API to assist in accurate image recognition; The network structure of the YOLOv8 target detection algorithm includes three parts: Backbone, Neck and Head; the Backbone part is a feature extraction network that extracts useful information from the input image, and adopts a fast-implemented CSP bottleneck module, and improves the feature extraction capability through the bottleneck block Bottleneck Block and the fast spatial pyramid pooling layer; the Neck part is located between the Backbone and the Head, and adopts a path aggregation network-feature aggregation network structure, and performs multi-scale feature fusion through the path aggregation network and the feature aggregation network; the Head part is responsible for predicting the category of the target, and adopts a decoupled head structure without an anchor frame to separate the classification and detection heads; The YOLOv8 target detection algorithm adds a CBAM attention mechanism to the Neck part. For the input image features of different scales extracted by its CSP LayerDarknet backbone network, the CBAM attention mechanism is used to perform attention calculations on the features in the channel and spatial dimensions, distinguish the features between different channels and capture the spatial structure in the image, and pass the image features of different scales processed by the CBAM attention mechanism to the Neck part for feature fusion; When the confidence level of the result obtained by the target detection algorithm is lower than the set threshold, the visual big model API is called to transmit the image data acquired by the camera in real time to the visual big model for recognition, and the object category and size information label is given.
3. The complex environment recognition and intelligent interaction system for blind navigation according to claim 1, characterized in that: In step S1, when a mark point appears within the positioning range of the client device, the map navigation module of the client device will generate a location information tag of the mark point.
4. The complex environment recognition and intelligent interaction system for blind navigation according to claim 1, characterized in that: In step S2, the large language model summarizes the information tags and outputs the text content in the form of voice to form an interaction with the user. The background server is mounted on the NVIDIANeMo framework. The specified format specifies the field information when the client device and the background server transmit data. The original binary data is Base64 encoded, and then the data is serialized into JSON format. The hash value is calculated to ensure integrity, and then the whole containing the data and the hash value is compressed, including the "image_data, image_info, audio_data, landmark_name" fields, where image_data, image_info, audio_data, and landmark_name represent image data, image information, voice information, and marker information, respectively.
5. The complex environment recognition and intelligent interaction system for blind navigation according to claim 1, characterized in that: In step S3, an autocorrelation analysis algorithm is used to reduce the influence of conditions that may affect acceleration changes during the travel process on the calculation of the step value; First, the three-axis acceleration α of the client device x 、a y , α z Take the modulus and calculate the overall acceleration α, and judge the user's motion state based on the standard deviation σ of the overall acceleration: Where u is the overall acceleration sequence of the user's step cycle {α1,α2,...,α k ,...,α N }, α k represents the kth overall acceleration, N represents the number of overall accelerations, and the user state is determined to be idle or walking according to the value of the standard deviation of the overall acceleration. When the standard deviation of the overall acceleration is less than the set threshold, it can be judged that the user's activity amplitude at this moment is small, and the movement state at this moment is determined to be idle, otherwise it is walking; Since the user walks continuously and rhythmically, the overall acceleration has obvious periodic changes over time. This periodic change can be used to calculate the autocorrelation between the overall acceleration of the current step cycle and the previous step cycle, and the user's motion state can be further judged based on the size of the autocorrelation. When the user walks, the acceleration sensor will continuously record data. After calculating the overall acceleration, the autocorrelation coefficient of the overall acceleration x(m,t) is calculated using the following formula, where m is the step cycle and t is the sampling period: Wherein, u(m,t) and σ(m,t) represent the mean and standard deviation of the acceleration sequence {a(p), a(p+1), ..., a(p+1-t}) in the current step cycle m within the sampling period t; u(m+t,t) and σ(m+t,t) are the values of u(m,t) and σ(m,t) after a sampling period t. p is an auxiliary time parameter, which can help calculate the appropriate cycle length from 0 to avoid calculation omissions. a(p) is the overall acceleration at time p, a(m+p) is the overall acceleration at time m+p, and a(m+p+t) is the overall acceleration at time m+p+t. When the sampling period t is close to the user's step period m, the walking states represented by a(m+p) and a(m+p+t) in the formula can be considered consistent, so the value of x(m,t) is close to 1. However, the step frequencies of different users or the same user at different times are inconsistent, so t is a variable. The framing algorithm needs to be used to dynamically determine t, that is, the value interval of t is selected as [t min , t max ], t min Represents the minimum value of the sampling period t, t max Represents the maximum value of the sampling period t, and calculates the new autocorrelation coefficient ρ(m,t): The t value when ρ(m, t) reaches its maximum value is the user's stable step cycle. In this cycle, the ratio of the walking distance to the pedometer step count is the average step length.
6. The complex environment recognition and intelligent interaction system for blind navigation according to claim 1, characterized in that: In step S4, the monocular vision technology uses a triangular transformation method to measure the distance between the object and the user. There are two cases: The first one is when the object is directly in front of the camera: Through inverse trigonometric functions, we can get the angle between the camera optical axis and the Y axis of the world coordinate system, and the angle between the line connecting the measured point and the origin of the image coordinate system and the camera optical axis. By subtracting the two angles, we can get the angle ∠Υ between the line connecting the measured point and the origin of the image coordinate system and the Y axis of the world coordinate system. By dividing the camera height H by the tangent value of ∠Υ, we can get the required distance in this case. The second case is when the object being measured is not directly in front of the camera: The point to be measured is recorded as Q, and its projection point in the image coordinate system is recorded as Q'. The distance between Q' and the camera is obtained from the first case, thereby obtaining the geometric distance from the origin of the camera coordinate system to the projection of the point to be measured on the Y axis of the world coordinate system. Since the triangle formed by the origin of the camera coordinate system, the point to be measured and the projection of the point to be measured on the Y axis of the world coordinate system has a similar relationship with the projection of the triangle in the image coordinate system, the required distance in this case can be calculated by measuring the distance from the origin in the camera coordinate system to the projection of the point to be measured, combined with the above similarity relationship.
Citation Information
Patent Citations
Blind person indoor navigation method based on inertial navigation and combined with Wi-Fi-assisted positioning
CN106248081A
Blind person intelligent navigation system and method based on image semantic segmentation
CN118298170A
Time-adjustable smartphone holder
KR1020220151850A