Orientation assistance system comprising means for acquiring a real or virtual visual environment, non-visual human-machine interface means and means for processing a digital representation of said visual environment
Patent Information
- Application Number
- JP2024545828
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-01
- Filing Date
- 2023-01-18
- Publication Date
- 2025-10-17
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the field of orientation assistance for visually impaired people or people moving in very poor visibility environments, for example firefighters moving in smoke-filled buildings or military personnel moving in darkness.
[0002] Various solutions are known, ranging from the aid of guide dogs, to marking the ground with guide strips, to installing audio beacons or using walking sticks to detect obstacles.
[0003] It has also been proposed to use haptic modes of information transmission, for example in the form of a connected wristband. Haptic technology uses the sense of touch to convey information. WearWorks offers a smart bracelet called the "Wayband" to guide the blind. The user starts by downloading an application on an associated smartphone and entering the desired address. The bracelet is linked to a GPS system and guides the user to their destination. If the user takes a wrong path, the bracelet vibrates; it stops vibrating when on the correct track. Tactile language is more sensitive, more intuitive, less intrusive, and relieves the over-sensation of the hearing, i.e., visually impaired.
[0004] Under French Patent No. 3100636(B1), the Applicant himself has patented an orientation assistance system comprising means for acquiring a real or virtual visual environment, non-visual human-machine interface means and means for processing a digital representation of said visual environment to provide electrical signals for controlling an interface consisting of a single tactile zone with a surface area between 60x60 mm and 150x150 mm and a bracelet with a set of NxM active spikes, N being between 5 and 100 and M being between 10 and 100, said digital representation processing means being configured to periodically extract at least one pulsed digital activation pattern for a subset of the spikes of said tactile zone.
[0005] Active belts have also been proposed to increase the surface area of the tactile zone.
[0006] prior art US Patent Publication No. 2013201308 describes the following steps: (1) capturing a black and white image to reduce detail and refine the image to obtain an object profile signal, and extracting profile information from the black and white image; (2) A visual impairment guidance method including: converting an object profile signal into a serial signal according to ergonomic characteristics, and transmitting the serial signal to an image sensory instrument, which converts the serial signal into a mechanical tactile signal to generate a feeler pin stimulus. Regarding the speed of tactile sensation for vision, an intermittent picture touch mode is used. The feeler pin array allows the visually impaired person to touch the shape of the object.
[0007] Optionally, this document proposes to explore position information of an object, and process the position information to obtain and prompt the distance and safe avoidance direction of the object. The explored position information from the object is processed to obtain and prompt the distance and safe avoidance direction of the object, so that the blind person can not only perceive the shape of the object, but also know the distance of the object.
[0008] US2019332175 relates to a wearable electronic tactile vision device configured to be attached to or worn by a user. The wearable electronic tactile vision device is configured to provide tactile feedback using pressurized air on the user's skin based on objects detected in the user's environment. Information about objects detected in the surroundings is captured using a digital camera, radar and / or sonar, and / or a 3D capture device such as a 3D scanner or a 3D camera attached to the wearable electronic tactile vision device. The wearable electronic tactile vision device is in the form of a helmet with at least two cameras positioned at the user's eye positions, or in the form of a T-shirt or other wearable accessory. [Background technology]
[0009] Both prior art documents propose to provide the user with a tactile translation that corresponds to an optical image obtained from a perspective view.
[0010] This is, of course, an obvious technique designed to compensate for the degradation of one of the senses, vision, by restoring the same information perceptible by another, touch.
[0011] The problem is that perception of the environment is not limited to "reading" a flat photographic image, but is the result of a complex process involving interpretation by the brain, which can provide a wealth of information, including depth, even when binocular vision is impaired.
[0012] Transposing the image into a tactile form does not allow the brain to benefit from this processing, resulting in confused and incomprehensible sensations that are largely excessive and useless. Summary of the Invention
[0013] Solution provided by the present invention In order to remedy the drawbacks of the prior art, the present invention, in its most general sense, relates to an orientation assistance device exhibiting the technical features set out in claim 1.
[0014] The solution provided by the present invention is not to transpose an optical image into a tactile image, but to generate, through a "depth scan" of the environment, a series of slice planes from a given image whose active pixels correspond to obstacles in the activated planes, providing the user with information with very few spikes that are activated when the environment is free of obstacles.
[0015] The orientation assistance system comprises means for acquiring a real or virtual visual environment, non-visual human-machine interface means, and means for processing a digital representation of said visual environment to provide electrical signals for controlling a haptic interface, said digital representation processing means being configured to periodically extract at least one pulsed digital activation pattern for a subset of spikes of said haptic zone. The haptic interface consists of a waist belt having an active surface of N×M spikes whose movement is controlled by an actuator, preferably a solenoid, where N and M are integers greater than or equal to 10. For each acquisition of said visual environment, the processing means provides a sequence of P activation frames for said actuator, where P is an integer between 2 and 15, preferably between 5 and 10, each frame corresponding to a representation of the environment in an incremental depth plane.
[0016] Preferably, the environment acquisition means comprises a spectacle frame carrying one or two cameras.
[0017] The invention also relates to a method for processing a digital representation of a visual environment to control a haptic interface consisting of a waist belt with active surfaces of N×M actuators, N and M being integers greater than or equal to 10, characterized in that for each acquisition of said visual environment, a sequence of P activation frames for said actuators is calculated, where P is an integer between 2 and 15, preferably between 5 and 10, each of the frames corresponding to a representation of the environment in an incremental depth plane.
[0018] According to one variant, it comprises a step of calculating digital images of N and M tactile pixels in a direction offset at a height of 10 to 100 cm from the ground.
[0019] According to another variant, the method comprises a step of calculating, for each digital image, a sequence of P successive frames corresponding to incremental depth planes.
[0020] Preferably, said step of calculating digital images of the N and M haptic pixels comprises a process consisting of assigning to each haptic pixel a density value which corresponds to the highest density value of the visual voxels which correspond to said haptic pixel.
[0021] According to one variant, said step of calculating the digital images of the N and M tactile pixels comprises a process consisting of assigning a non-zero density value to the areas of the visual image that correspond to the holes.
[0022] According to one variant, said step of calculating digital images of N and M tactile pixels comprises a process consisting of assigning a non-zero density value to areas of the visual image that correspond to obstacles by an automatic recognition process.
[0023] According to a particular embodiment, the step of calculating the digital images of the N and M tactile pixels includes a process consisting of removing voxels outside the user's traffic lane before calculating the digital images of the N and M tactile pixels established from only the remaining voxels.
[0024] Preferably, the positions of the voxels are modified according to their depth in order to take full account of the display volume.
[0025] Preferably, the step of calculating digital images of the N and M tactile pixels comprises a process of reducing the processed voxels as a function of parameters including the user's movement speed and / or the movement speed of an object within the field of view of the visual acquisition means and / or the distance of the object, before calculating the digital images of the N and M tactile pixels established from only the remaining voxels.
[0026] In one variant, the method includes a calculation step for converting the distance to the camera into a distance to the user.
[0027] According to a particular embodiment, the processed image bursts are recalculated if the orientation of the viewing direction of the environment changes. [Brief description of the drawings]
[0028] The present invention will now be described in more detail with reference to non-limiting exemplary embodiments which identify the advantages and considerations set forth above. A more specific description of the present invention is provided briefly below.
[0029] [Figure 1] 1 shows a schematic diagram of a system according to the present invention. [Diagram 2] A diagram of a visual image is shown. [Diagram 3] A diagram of a tactile image is shown. [Figure 4] 1 shows a diagram of a sequence of haptic frames. [Diagram 5] 1 shows a schematic diagram of a vertical tunnel. [Figure 6] 1 shows a schematic diagram of the field of view as a function of movement speed. [Figure 7] A schematic diagram of a horizontal tunnel is shown. [Figure 8] 13 shows exemplary code for processing horizontal fields of view. [Figure 9]1 shows exemplary code for a tactile imaging application. [Figure 10] 1 shows exemplary code for an accelerated haptic imaging application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0030] A non-limiting example developed below comprises a spectacle frame (10) equipped with means for acquiring the environment, for example cameras (11, 12) used to acquire data about the environment in real time to provide digital images that control the action of a tactile transducer, which generates the action in the form of pressure on the skin by means of an electromagnetic or electromechanical actuator, or in the form of an electrical impulse or light.
[0031] It should be recalled that the system can also be used for augmented reality gaming or training applications, with images provided by a video source.
[0032] A computer processes the extraction of images from the sensor unit and then generates from them a 3D depth map. The computer transmits this map to a haptic device such as a grid of solenoids or spikes (small linear actuators that can be raised or lowered) integrated into a back belt (20). This belt (20) is equipped with a set of solenoids arranged on supports (21-24) so as to form a matrix of, for example, 20x40 pixels. These solenoids are arranged to form a regular matrix, preferably with a constant pitch. An electronic circuit receives visual signals and processes them to control the solenoids, generating sensations on the user's back that are easy to interpret after a learning period. The waist belt (20) can be worn on a lightweight fabric garment (shirt, polo shirt, blouse) or directly on the skin.
[0033] The surface of the active matrix formed by the solenoids covers an extended lumbar area for good resolution and comfort of use.
[0034] Rendering an image of the real environment into a haptic image consists in dividing the depth map calculated from the visual image into several successive layers, each of which determines a virtual or haptic image that controls the activation of the haptic device, so that the nearest objects are displayed first, followed by slightly more distant objects, and so on, until a maximum visual distance is reached (usually about 10 meters). This forms a kind of scanning of the environment, gradually disappearing and displaying what is encountered at every moment. This scanning results in a burst of virtual images lasting about 100 milliseconds, built up of about 10 haptic images corresponding to successive planes, before resuming with a new burst corresponding to a new environment resulting from the user's movement or a change in the orientation of the real image due to a change in the position of the head or the video image.
[0035] Visual Image Processing The cameras (11, 12) capture binocular images to reconstruct a digital image with depth information. The first step is to build a grayscale image. For each pixel of the optical image (100), the tactile image (200) is transferred in grayscale according to the distance of each point from the camera.
[0036] Depth information can also be determined using a single camera with appropriate image processing.
[0037] It should be noted that this tactile image (200) could also be calculated from digital images provided by a lidar without departing from the invention.
[0038] This haptic image (200) is then decomposed into a sequence of incremental haptic frames (301-307), each corresponding to a depth plane: the first haptic frame (301) of the sequence corresponds to the obstacle zone in the plane closest to the user, the second first haptic frame (301) of the sequence corresponds to the obstacle zone in the plane closest to the user, each subsequent haptic frame (302) corresponds to the obstacle zone in the next plane offset by one step from the previous haptic frame, e.g. a distance of 30 centimeters, and so on.
[0039] The grey scale of a haptic frame (301-307) encodes the type of action of the corresponding solenoid, for example the frequency of vibration or the duration of vibration during the activation time of the corresponding haptic frame.
[0040] Thereby, the haptic image (200) is replaced by a time-scan of haptic frames (301-307), which are integrated by the user to perceive a depth representation of the user's environment.
[0041] Other procedures are applied to improve the clarity of tactile perception. Correcting the visual image (100) to create a synthetic image based on observations along a horizontal axis lowered to the height of the user's feet several tens of centimeters above the ground. highlighting pixels corresponding to small obstacles in the visual image (100) so that they occupy at least one pixel in the tactile image (200); Reprocessing holes to enhance the grayscale of the corresponding areas on the tactile image (200). Grayscale enhancement on the tactile image (200) of a region of interest determined by automatic recognition, e.g., by supervised learning.
[0042] 1 illustrates an embodiment of a process for producing a tactile image (200).
[0043] The following variables are used to describe an embodiment of the process: M represents the haptic matrix applied to the user H represents the height (in pixels) of the haptic matrix M W represents the width of the haptic matrix M (in pixels) p is the precision level DM represents the depth map derived by the sensors (11, 12). dmH is the height of the restored depth map (in pixels) = p * H, and dmW is the width of the restored depth map (in pixels) = p * W and FOVv represents the vertical field of view (in degrees) of the camera used FOVh represents the horizontal field of view (in degrees) of the camera used hu represents the height at which the camera is placed (the height of the user photographed at eye level) DISTANCE_MAX is the maximum viewing distance set by the user SPEED represents the display speed set by the user. PAUSE_TIME represents the pause time (set by the user) between the display of two images. MAT[x,y] corresponds to the value of the matrix MAT at the coordinate [x,y] MAT[x] corresponds to the x-th column of the matrix MAT, and LIST[x] corresponds to the x-th value of the list LIST.
[0044] x=f(arguments) means that the value of x is a function (proportional relationship) of one or more arguments
[0045] / / means that what follows is a comment
[0046] Step 1: Get the depth map. This first step involves the computation of a 3D image of size dmW from two images acquired by cameras (11) and (12), or by a lidar, or by a binocular virtual image source. * This consists of computing a depth map of dmH.
[0047] The process has a well-known form and generally involves the following steps. a. Acquiring two images of the same scene simultaneously by two cameras (11, 12) with a known separation, or acquiring successive images from the same sensor. b. A calibration step consisting of determining the intrinsic and extrinsic parameters of the geometric model of the stereo sensor. c. A pixel matching step to find pairs of pixels on the two images that correspond to the projections of the same scene elements. d. A 3D reconstruction step consisting of calculating, for each pixel, the position in space of the point projected onto this pixel.
[0048] The result of this first step is a size dmW * A visual image (100) of dmH, where each point is a voxel defined by the coordinates of each point in space, where the origin is at the user's head, the x and y axes are perpendicular to the camera's line of sight, and the z coordinate is the distance from the user's head.
[0049] The location of the voxels is adapted to the depth in order to take full advantage of the tactile display capabilities.
[0050] Reduced matrix resolution without losing important information The objective of this process is to reduce the size of the depth map derived by the sensor to the size of the hip representation (spike matrix). The problem with traditional resolution reduction is that some information may be lost. For example, if a very narrow pole is in front of the user, it may not be displayed, which is a major safety issue. To remedy this problem, the following resolution reduction algorithm is used, which has the advantage of retaining the closest (and therefore most important) objects in each zone. This algorithm is based on a resolution of dmW * It takes as input a depth map of dmH and has resolution W * Returns the H' matrix.
[0051] This can be done by the following code: FOR i from 0 to H(excluded) FOR j from 0 to W(excluded) Create an empty list FOR k from 0 to p(excluded) FOR l from 0 to p(excluded) Add DM[p * i+k,p * j+l] to the list END of loop END of loop k Sort list by ascending values M[j,i]=first quartile of list(=list[list size / 4]) END of loop j END of loop i
[0052] Ground hole detection processing It is not easy for a user to find a pothole by perceiving the absence of activation of a certain solenoid, since the absence of information is very difficult to perceive. One solution is to identify the pothole and modify the grayscale of the corresponding pixel in the tactile image. Such processing can be performed by a program whose algorithms identify the location of the pothole and transmit its location to the display, as described in detail below. Create a list of integers of size W:holes Let m be the margin of error(at least 20%) FOR i from 0 to W(excluded) Create variable dMax=(hu / sin(FOVv / 2)) IF M[i,H-1]>dMax* (1+margin / 100) THEN holes[i]=1 ELSE holes[i]=0 END of loop i
[0053] Transmission of holes to the display Similarly, thanks to artificial intelligence and image processing, it is possible to generate a "pit" list of the locations of obstacles (roots, narrow walkways, etc.) that are too small to display.
[0054] Distance change Another process involves converting the distance to the camera into the distance to the user.
[0055] Human distance perception is based on the whole body, not the eyes. This conversion is usually performed intuitively by the brain. In the context of the present invention, this correction is performed upstream to simplify the process of skin perception, for example through algorithmic processing such as:
[0056] FOR i from H / 2 to H(excluded) Create variable
[0057]
number
[0058] FOR j from 0 to W(excluded) M[j,i] = cos(α)M[j,i] END of loop j END of loop i
[0059] "Tunnel vision" processing Our eyes perceive everything that is visible in our field of vision, but our brains do not process all the information. This is the difference between seeing and seeing.
[0060] To avoid information overload, the present invention provides a process that limits the information to only those obstacles present in the "virtual corridor" (30) ahead of the user, eliminating the most useless information (31, 32) (see FIG. 5), and then re-arranges the truncated matrix so that the information retained uses the entire display matrix.
[0061] This process is carried out by a program corresponding to the following algorithm: / / replace the values to be deleted with a constant, “CODE” Create a constant integer CODE FOR i from 0 to H (excluded) Create variable
[0062]
number
[0063] IF i < H / 2 then Create variable ht FOR j from 0 to W ht = M[j,i] * sin(α) IF ht>(hv-hu) THEN M[j,i]=CODE END of loop j END of IF ELSE Create variable hb FOR j from 0 to W hb = M[j,i] * sin(α) IF hb>hu THEN M[j,i]=CODE END of loop j END of ELSE END of loop i / / then “stretch” the matrix column by column to fill in deleted values FOR i from 0 to W Create variable-size list:column=M[i] Remove elements from column for which:value==CODE Create double variable:r=size(column) / W Create list with size W:newColumn FOR j from 0 to H newColumn[j]=column[floor(r * j)] / / floor=rounded down END of loop j M[i] = newColumn END of loop i
[0064] "Horizontal field of view" processing As with the vertical axis detailed above, a full horizontal field of view is not always useful and can lead to information overload. However, certain information must not be lost, which is why no passage is defined here as before. Here, the visual reduction depends on several parameters that are targeted in the algorithm. User Speed object velocity object distance
[0065] Figure 6 shows a top view. The higher the user's speed v, the smaller the field of view. Figure 7 shows the exclusion zones (41, 42) and the retention zone (40).
[0066] The processing algorithm is illustrated in FIG.
[0067] Matrix Application Processing In this display, everything depends on the distance to the nearest object in the field of view. Three values are derived from it: maxTime: defines the time it should take to display the entire image (if the time has elapsed, the display moves to the next image), allowing the user to intuitively know whether they are close to an obstacle or not maxDistance: defines the maximum viewing distance. The algorithm will focus on the nearest object to avoid information overload. For example, if the nearest obstacle is 1m from the user, the program will only display obstacles that are between 1m and 2.5m away. distanceByLayer: To display an image in 3D, the algorithm very quickly displays successive 2D layers. For example, it displays a first layer with objects from 0cm to 30cm, then objects from 30cm to 60cm, and so on. To make this distance consistent with the context, it evolves proportionally to the closest one.
[0068] FIG. 9 shows an example implementation of this algorithm.
[0069] indication For users who wish to do so, a faster but more learning-requiring display mode is available.
[0070] In this mode, the algorithm in charge of the display only updates the changes in the current matrix compared to the previous one. Thus, if everything is static, nothing is displayed, but as soon as an object or the user moves, the user will see a change. This process is shown in Figure 10.
[0071] Motion blur The matrix sent to the user may update before being completely displayed and continue to be displayed. The bursts are displayed in approximately 100 milliseconds. If the user rotates his or her head while the sequence is being applied, one variation is to recalculate the virtual image and apply the modified burst from the new camera orientation.
[0072] Customize your settings The tactile sensitivity of the dorsal zone and the clarity of the tactile stimulation vary from person to person. To make it easier for the user to learn and understand this guidance mode, the invention optionally provides a configuration layer that optimizes the adaptation to a particular user. This configuration software layer consists in determining the way in which the real images are transformed into virtual images corresponding to the depth layer, in particular the periodicity of the bursts, the duration of the tactile application of each virtual depth image, the possibility of introducing virtual images at the beginning and / or end of the bursts, the resolution of the virtual images, etc.
[0073] These parameters can be defined by a supervised learning process, using reference paths and taking into account the user's error types.
Claims
1. 1. An orientation assistance system comprising: means for acquiring a real or virtual visual environment; a non-visual human-machine interface means; means for processing the digital representation of the visual environment to provide electrical signals for controlling a haptic interface; Equipped with the means for processing the digital representation periodically extracts at least one pulsed digital activation pattern for a subset of spikes in the tactile zone; the haptic interface consists of a waist belt with an active surface of N x M spikes, where N and M are integers greater than or equal to 10; - said means for processing provides a sequence of P activation frames for said spike, where P is an integer between 2 and 15, preferably between 5 and 10, each frame corresponding to a representation of the environment in an incremental depth plane; An orientation assistance system comprising:
2. The spike is activated by a solenoid. An orientation assistance system comprising means for acquiring a real or virtual visual environment according to claim 1.
3. characterised in that the means for acquiring the visual environment are constituted by a spectacle frame (10) carrying one or two cameras (11, 12), An orientation assistance system comprising means for acquiring a real or virtual visual environment according to claim 1.
4. 1. A method for processing a digital representation of a visual environment to control a haptic interface consisting of a waist belt having active surfaces of N×M actuators, where N and M are integers greater than or equal to 10, characterized in that for each acquisition of said visual environment, a sequence of P activation frames for spikes is calculated, where P is an integer between 2 and 15, preferably between 5 and 10, each of said frames corresponding to a representation of the environment in an incremental depth plane.
5. Calculating digital images of N and M tactile pixels in an offset direction at a level of 10-100 cm from the ground, 5. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 4.
6. for each digital image, calculating a sequence of P consecutive frames corresponding to incremental depth planes, A method for processing a digital representation of a visual environment to control a haptic interface according to claim 4 or 5.
7. said step of calculating the digital images of the N and M haptic pixels comprises a process consisting of assigning to each haptic pixel a density value corresponding to the highest density value of the visual voxels corresponding to said haptic pixel; A method for processing a digital representation of a visual environment for controlling a haptic interface according to claim 5.
8. said step of calculating digital images of the N and M tactile pixels comprises a process of assigning non-zero density values to areas of the visual image corresponding to holes; 6. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 5.
9. said step of calculating the digital images of the N and M tactile pixels comprises a process consisting of assigning a non-zero density value to areas of the visual image corresponding to obstacles by an automatic recognition process; 6. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 5.
10. said step of calculating the digital images of the N and M haptic pixels comprises a process consisting of removing voxels outside the user's lane before calculating said digital images of the N and M haptic pixels established only from the remaining voxels, A method for processing a digital representation of a visual environment for controlling a haptic interface according to claim 5.
11. said step of calculating the digital images of N and M haptic pixels comprises a process consisting of reducing said processed voxels as a function of parameters including the speed of movement of the user and / or the speed of movement of an object in the field of view of the means for acquiring the visual environment and / or the distance of said object, before calculating said digital images of N and M haptic pixels established only from the remaining voxels; A method for processing a digital representation of a visual environment for controlling a haptic interface according to claim 5.
12. a calculation step for converting a distance relative to the camera into a distance relative to the user, 5. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 4.
13. the processed image bursts are recalculated when the orientation of the viewing direction of the environment changes; 5. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 4.
14. the positions of the voxels are modified according to their depth in order to take into account the display volume; 5. A method for processing a digital representation of a visual environment to control a haptic interface according to claim 4.