Artificial eye space perception method based on binocular camera
The video stream is obtained through binocular cameras, distortion correction and object detection are performed, the distance between the target and the user is calculated using the parallax map, and the distance between the obstacle and the user is described by the perception model, the problems of messy information and scene applicability in the prior art are solved, and effective spatial perception of user travel and daily life scenarios are realized.
Patent Information
- Application Number
- CN202510135018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art uses traditional distance-seeking methods, resulting in messy information in binocular camera object detection, interfering with user judgment, and is only suitable for guided scenes and cannot meet the needs of daily life scenes.
The video stream is obtained through binocular cameras, distortion correction and object detection are performed, the distance between the target and the user is calculated using the parallax map, and the distance between the obstacle and the user is described through the perception model, so as to realize spatial perception of the user's travel and daily life scenes.
It realizes effective spatial perception of users' travel and daily life scenarios, reduces information mess, enhances users' judgment ability, and is still feasible in an outage environment, ensuring data security and personal privacy.
Smart Images

Figure CN120070273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to an artificial eye space perception method based on a binocular camera. Background Art
[0002] With the progress of technology and the gradual development of the prosthetics field, more and more people have paid attention to the needs of blind artificial eyes. However, due to the difficulty in training guide dogs, a large number of blind people still rely on traditional guide canes to travel, which poses a great potential safety hazard. At the same time, there is no product like an artificial eye for the blind on the market at present.
[0003] The existing invention patent with the publication number of CN118570558A discloses a binocular target detection method and system for a guide blind glasses. The invention uses the images obtained by a binocular camera, analyzes the angular difference data, performs hierarchical analysis on the depth information and reconstructs a three-dimensional structure to provide spatial perception for visually impaired users, and classifies the depth information by distance and assigns corresponding sound and vibration intensities to be transmitted to the users.
[0004] Although the above method introduces the method of binocular camera target detection, it uses a traditional distance calculation method, which will return all detected targets to the user at the same time, resulting in messy information and interfering with the user's judgment. In addition, this method is only applicable to the guide blind scenario when the user travels and is not applicable in daily life scenarios. To solve these problems, the present invention proposes an artificial eye system and method based on a binocular camera. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an artificial eye space perception method based on a binocular camera, which solves the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0007] The artificial eye space perception method based on a binocular camera includes the following steps:
[0008] S1. Distortion correction is performed on the images captured by the binocular camera and a video stream is generated;
[0009] S2. Artificial eye space perception:
[0010] A target detection model performs target detection on the images of the video stream to determine each target in the image;
[0011] Based on the view difference between the left-eye image and the right-eye image, the distance between the target and the user is calculated;
[0012] The image after target detection is transformed into an image with detection marks;
[0013] S3. The user obtains spatial information:
[0014] The perception large model identifies the image with the detection mark and describes to the user the content of the image with the detection mark and the distance between the obstacle and the user.
[0015] The perception large model interacts with the user and fulfills the user's further description requirements for the content of the image with the detection mark.
[0016] Furthermore, the specific steps of S1 include:
[0017] S1.1. Calibrate the parameters of the binocular camera to be consistent.
[0018] S1.2. Calibrate the binocular camera and save the calibration parameters as a configuration file.
[0019] S1.3. Collect the real-time images captured by the left-eye camera and the right-eye camera at a preset frame rate, and correct the distortion of the collected real-time images using the configuration file saved in S1.2.
[0020] S1.4. Save the corrected images in step S1.3 as a video stream.
[0021] S1.5. Analyze the video stream frame by frame and store it in the left-eye image folder and the right-eye image folder respectively.
[0022] Furthermore, the specific steps of S1.5 are:
[0023] Analyze the video streams of the left-eye camera and the right-eye camera in parallel at the same frame rate.
[0024] Unzip the video stream of the left-eye camera into left-eye images according to the naming method of timestamp plus L, and store them in the left-eye image folder. At the same time, unzip the video stream of the right-eye camera into right-eye images according to the naming method of timestamp plus R, and store them in the right-eye image folder.
[0025] Furthermore, in S1.5, when the left-eye video stream and the right-eye video stream captured at the same moment are unzipped into frames of images, the timestamps of the left-eye images are the same as those of the right-eye images.
[0026] Furthermore, the specific steps of S2 are:
[0027] Obtain the left-eye images in the left-eye image folder that have not participated in object detection in the order of timestamps.
[0028] Use the object detection model to detect the left-eye images, make predictions on the left-eye images, and return the prediction results.
[0029] Extract the right-eye image with the same timestamp as the left-eye image, and calculate the disparity map using the extracted left-eye image and right-eye image;
[0030] Using the prediction result and the disparity map, calculate the distance between the detected target and the user;
[0031] Draw the target detection box on each target in the left-eye image, and return the image with detection marks;
[0032] Delete the left-eye image and right-eye image that have participated in target detection to save memory.
[0033] Further, the distance between the target and the user is marked in the upper left corner of the target detection box.
[0034] Further, the content describing the image with detection marks is completed by the image-to-text module of the perception large model, and the distance between the obstacle and the user is completed by the attention mechanism of the knowledge large model.
[0035] Further, the specific steps for the perception large model to interact with the user and complete the user's further description requirements for the content of the image with detection marks are as follows:
[0036] The perception large model interacts with the user and determines whether the user needs further description:
[0037] Yes, the perception large model describes each target in detail according to the user's needs;
[0038] No, exit the current step, continue to execute S2, and repeat S2 - S3 using the next frame of image.
[0039] Further, the image-to-text module includes a local image-to-text module and a server image-to-text module; when the network speed meets the threshold or the user needs to describe each target in detail, the server image-to-text module works; when it is necessary to prevent leakage or the network speed does not meet the threshold, the local image-to-text module works.
[0040] Further, the target detection model can be optimized according to the hardware performance of the deployment platform, and the optimization operations include pruning, distillation, and quantization.
[0041] The present invention provides a method for spatial perception of an artificial eye based on a binocular camera. Compared with the prior art, it has the following beneficial effects: obtaining a video stream through a binocular camera; parsing the obtained video stream and decompressing the video stream of the binocular camera into synchronized frame images; using a target detection model to predict the image of the left-eye camera, and using the synchronized frame images of the binocular camera to obtain the distance between the detected target and the user; using an image-to-text large model to describe the content in the image, and using the attention mechanism in the large model to focus on describing the distance between the obstacle and the user, and using the large model to interact with the user to complete the user's needs, achieving the effect of taking care of the user's travel and daily life scenarios. At the same time, considering the issues of data security and personal privacy, even in an environment without network connection, the system can still operate normally. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0043] Figure 1 The schematic flowchart of the method for spatial perception of an artificial eye based on a binocular camera according to the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0045] To solve the technical problems in the background art, the following method for spatial perception of an artificial eye based on a binocular camera is given:
[0046] Combined with Figure 1 As shown, the method for spatial perception of an artificial eye based on a binocular camera provided by the present invention includes S1, performing distortion correction on the images captured by the binocular camera and generating a video stream;
[0047] S2, artificial eye spatial perception:
[0048] The target detection model performs target detection on the images of the video stream to determine each target in the image;
[0049] Based on the view difference between the left-eye image and the right-eye image, the distance between the target and the user is obtained;
[0050] The image after object detection is transformed into an image with detection marks;
[0051] S3. The user obtains spatial information:
[0052] The perception large model recognizes the image with detection marks and describes the content of the image with detection marks and the distance between the obstacle and the user to the user;
[0053] The perception large model interacts with the user and completes the user's further description requirements for the content of the image with detection marks.
[0054] As an embodiment, the specific steps of S1 include:
[0055] S1.1. Calibrate the parameters of the binocular camera to be consistent;
[0056] S1.2. Calibrate the binocular camera and save the calibration parameters as a configuration file;
[0057] S1.3. Collect the real-time images captured by the left-eye camera and the right-eye camera at a preset frame rate, and use the configuration file saved in S1.2 to correct the distortion of the collected real-time images;
[0058] S1.4. Save the corrected images in step S1.3 as a video stream;
[0059] S1.5. Analyze the video stream frame by frame and classify and store it in the left-eye image folder and the right-eye image folder.
[0060] As an embodiment, the specific steps of S1.5 are:
[0061] Parallelly analyze the video streams of the left-eye camera and the right-eye camera at the same frame rate;
[0062] Decompress the video stream of the left-eye camera into left-eye images according to the naming method of timestamp plus L and store them in the left-eye image folder. At the same time, decompress the video stream of the right-eye camera into right-eye images according to the naming method of timestamp plus R and store them in the right-eye image folder. The video stream files in the right-eye video stream folder and the left-eye video stream folder that exceed the set threshold are identified as expired data, and the expired data is cleared.
[0063] As an embodiment, in S1.5, when the left-eye video stream and the right-eye video stream captured at the same moment are decompressed into frames of images, the timestamps of the left-eye images are the same as those of the right-eye images.
[0064] As an embodiment, the specific steps of S2 are as follows: Obtain the left-eye images in the left-eye image folder that have not participated in object detection in the order of timestamps; Use the object detection model to detect the left-eye images, make predictions on the left-eye images, and return the prediction results; Take out the right-eye images with the same timestamps as the left-eye images, and use the extracted left-eye images and right-eye images to calculate the disparity map; Use the prediction results and the disparity map to calculate the distance between the detected objects and the user; Draw the object detection frames on each object in the left-eye image, and return the image with detection marks; Delete the left-eye images and right-eye images that have participated in object detection to save memory.
[0065] As an embodiment, the distance between the object and the user is marked in the upper left corner of the object detection frame.
[0066] As an embodiment, the content describing the image with detection marks is completed by the image-to-text module of the perception large model, and the distance between the obstacle and the user is completed by the attention mechanism of the knowledge large model.
[0067] As an embodiment, the specific steps for the perception large model to interact with the user and complete the user's further description requirements for the content of the image with detection marks are as follows:
[0068] The perception large model interacts with the user and determines whether the user needs further description:
[0069] If yes, the perception large model describes each object in detail according to the user's needs;
[0070] If no, exit the current step, continue to execute S2, and repeat S2 - S3 using the next frame of image.
[0071] As an embodiment, the image-to-text module includes a local image-to-text module and a server image-to-text module; When the network speed meets the threshold or the user needs to describe each object in detail, the server image-to-text module works; When it is necessary to prevent leakage of secrets or the network speed does not meet the threshold, the local image-to-text module works.
[0072] As an embodiment, the object detection model can perform optimization operations according to the hardware performance of the deployment platform, and the optimization operations include pruning, distillation, and quantization.
[0073] The specific implementation process of the above technical solution is as follows:
[0074] 1. Set the shooting parameters of the left and right binocular cameras to be the same, including: image format and image size. The video stream of the left camera is named according to the timestamp plus L and saved in the video stream folder corresponding to the left camera. The video stream of the right camera is named according to the timestamp plus R and saved in the video stream folder corresponding to the right camera. At the same time, the video stream files in the left and right video stream folders that exceed the set threshold are identified as expired data and the expired data is cleared;
[0075] 2. Analyze the video stream and decompress the video stream into frames of images. According to the hardware performance or usage scenario, the video stream can be parsed into images at the set frame rate. When the memory is insufficient, downsampling is used to reduce the frame rate of converting the video stream into image data, so as to reduce the pressure on memory resources. When decompressing the video stream into image data, in order to ensure the accuracy of the subsequent distance calculation, it is necessary to ensure that the decompressed images of the left and right cameras are in a frame-synchronized state, that is, for any frame of the image of the left camera after decompression, the corresponding image of the right camera has the same shooting time;
[0076] 3. Use the object detection model to perform object detection on the parsed image obtained by any camera. In this embodiment, the object detection is performed on the image of the left camera, and the distance between the detected object and the camera is calculated using the images of the left and right eyes. Since this system is an artificial eye designed for the blind, theoretically the distance between the object and the camera is the distance between the object and the user. The distance here includes the Euclidean distance and the distances in the horizontal and vertical directions;
[0077] 4. Select the large model for text generation from images according to the user's needs. If the network speed meets the threshold requirements and a large model with accurate description is to be selected, the system selects the large model deployed on the server side for image description. If in a relatively private scenario, the user is worried about privacy leakage, or when the network speed does not meet the threshold requirements, the system automatically selects the large model deployed locally to describe the content in the image, and uses the attention mechanism in the large model to focus on describing the distance between the obstacle and the user. When the user is in a moving state, the distance to the nearest obstacle to the user is preferentially described, and at the same time, the large model is used to interact with the user to complete the user's needs.
[0078] The present invention obtains a video stream through a binocular camera; analyzes the obtained video stream, decompresses the video stream of the binocular camera into synchronous frame images; uses an object detection model to predict the image of the left-eye camera, and uses the synchronous frame images of the binocular camera to obtain the distance between the detected object and the user; uses a text-to-image large model to describe the content in the image, and uses the attention mechanism in the large model to focus on describing the distance between the obstacle and the user, and interacts with the user using the large model to complete the user's requirements, achieving the effect of taking care of the user's travel and daily life scenarios. At the same time, considering the issues of data security and personal privacy, the system can still operate normally even in an environment without network connection.
[0079] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.
[0080] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0081] In the several embodiments provided in this article, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be in electrical, mechanical, or other forms of connection.
[0082] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments in this article.
[0083] In addition, in each of the embodiments described herein, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0084] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution described herein, or the part that contributes to the prior art, or all or part of the technical solution may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments herein. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0085] Specific embodiments are used in this article to elaborate on the principles and implementation manners of this article. The description of the above embodiments is only used to help understand the method and its core idea of this article; at the same time, for those of ordinary skill in the art, according to the idea of this article, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this article.
Claims
1. A method for artificial eye spatial perception based on binocular camera, characterized in that: The steps include: S1, performing distortion correction on the image captured by the binocular camera and generating a video stream; S2. Prosthetic Eye Spatial Perception: The target detection model performs target detection on the images in the video stream and determines each target in the image; Calculate the distance between the target and the user based on the view difference between the left-eye image and the right-eye image; The image after target detection is transformed into an image with detection mark; S3. User obtains spatial information: The perception model recognizes the image with the detection mark and describes the content of the image with the detection mark and the distance between the obstacle and the user to the user; The perception model interacts with the user and completes the user's needs for further description of the content of the image with the detection mark.
2. The method for spatial perception of artificial eyes based on binocular cameras according to claim 1, characterized in that: The specific steps of S1 include: S1.
1. Calibrate the binocular camera parameters to make them consistent; S1.2, calibrate the binocular camera and save the calibration parameters as a configuration file; S1.3, collecting real-time images captured by the left camera and the right camera at a preset frame rate, and using the configuration file saved in S1.2 to perform distortion correction on the collected real-time images; S1.4, saving the image corrected in step S1.3 as a video stream; S1.
5. Analyze the video stream frame by frame and store them in the left-eye image folder and the right-eye image folder.
3. The method for spatial perception of artificial eyes based on binocular cameras according to claim 2, characterized in that: The specific steps of S1.5 are: Parse the video streams of the left and right cameras in parallel at the same frame rate; The video stream of the left camera is decompressed into a left-eye image according to the naming method of timestamp plus L, and stored in the left-eye image folder. At the same time, the video stream of the right camera is decompressed into a right-eye image according to the naming method of timestamp plus R, and stored in the right-eye image folder.
4. The method for spatial perception of artificial eyes based on binocular cameras according to claim 3, characterized in that: In S1.5, when the left-eye video stream and the right-eye video stream shot at the same time are decompressed into frame images, the timestamp of the left-eye image is the same as the timestamp of the right-eye image.
5. The method for spatial perception of artificial eyes based on binocular cameras according to claim 3, characterized in that: The specific steps of S2 are: Obtain the left-eye images that have not yet participated in target detection in the left-eye image folder in the order of timestamps; Use the target detection model to detect the left eye image, make predictions on the left eye image, and return the prediction results; Take out the right eye image with the same timestamp as the left eye image, and use the extracted left eye image and right eye image to calculate the disparity map; Using the prediction result and the disparity map, calculating the distance between the detected target and the user; Draw the target detection box on each target in the left image, and return the image with the detection mark; Delete the left and right images that have participated in target detection to save memory.
6. The method for spatial perception of artificial eyes based on binocular cameras according to claim 5, characterized in that: The distance between the target and the user is marked in the upper left corner of the target detection frame.
7. The method for spatial perception of artificial eyes based on binocular cameras according to claim 1, characterized in that: The description of the content of the image with the detection mark is completed by the image-to-text module of the perception model, and the description of the distance between the obstacle and the user is completed by the attention mechanism of the knowledge model.
8. The method for spatial perception of artificial eyes based on binocular cameras according to claim 7, characterized in that: The perception model interacts with the user and completes the user's further description of the content of the image with the detection mark. The specific steps are: The perception model interacts with the user and determines whether the user needs further description: If yes, the perception model describes each target in detail according to user needs; If not, exit the current step, continue to execute S2, and repeat S2-S3 using the next frame image.
9. The method for spatial perception of artificial eyes based on binocular cameras according to claim 8, characterized in that: The image-to-text module includes a local image-to-text module and a server image-to-text module; when the network speed meets the threshold or the user needs to describe each target in detail, the server image-to-text module works; when it is necessary to prevent leakage or the network speed does not meet the threshold, the local image-to-text module works.
10. The method for spatial perception of artificial eyes based on binocular cameras according to claim 1, characterized in that: The target detection model can be optimized according to the hardware performance of the deployment platform, and the optimization operations include pruning, distillation, and quantization.
Citation Information
Patent Citations
Binocular target detection method and system for blind guiding glasses
CN118570558A