Mobile, mobile control system, and mobile control method
The mobile body control system uses structured and normal lighting alternation to achieve dense environmental measurement and accurate self-position estimation in nuclear reactor environments, addressing the limitations of radiation-hardened cameras and patterned lighting noise.
Patent Information
- Application Number
- JP2024039259
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-29
AI Technical Summary
Existing environmental recognition technologies for nuclear reactor environments face challenges in achieving both dense environmental measurement and highly accurate self-position estimation due to the limitations of radiation-hardened cameras and the noise introduced by patterned lighting in feature-based vSLAM systems.
A mobile body equipped with an image sensor capable of acquiring stereo information, an illuminator for structured and normal lighting, and a control system that alternates between structured and normal lighting to determine feature points and stereo disparity, creating an environmental map using a combination of these methods.
This approach enables both dense environment measurement and highly accurate self-location estimation, overcoming the limitations of radiation-hardened cameras and patterned lighting noise.
Smart Images

Figure 2025140086000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a mobile object, a mobile object control system, and a mobile object control method. [Background technology]
[0002] Decommissioning work at nuclear power plants requires remotely operated survey robots and work robots for removing debris, as they are carried out in high-radiation environments. To operate such robots, environmental recognition technology equipped with sensors such as radiation-resistant cameras that can operate in high-radiation environments is required.
[0003] In environmental recognition technology for use inside such reactor buildings, work robots move around, so recognition is required in 3D at actual size. Radiation-resistant (radiation-hardened) cameras are available as sensors for recognizing the real environment in extreme environments with high radiation concentrations, but radiation-hardened cameras are large, making it difficult to mount multiple cameras, such as general stereo cameras, on a robot large enough to move around inside a decommissioned reactor. Furthermore, because cameras are prone to malfunction in a radiation environment, it is not desirable to apply a system that relies on multiple cameras, such as stereo cameras.
[0004] In response to this, Patent Document 1 proposes that, for the purpose of stably controlling the movement of a moving body, "an information processing device characterized by comprising: an input means for receiving input of image information acquired by an imaging unit mounted on a moving body, each light receiving unit on an imaging element of which is composed of two or more light receiving elements; a holding means for holding map information; an acquisition means for acquiring the position and orientation of the imaging unit based on the image information and the map information; and a control means for obtaining a control value for controlling the movement of the moving body based on the position and orientation acquired by the acquisition means." [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-125345 Summary of the Invention [Problem to be solved by the invention]
[0006] According to Patent Document 1, the object can be achieved.
[0007] However, when it comes to environmental recognition inside a nuclear reactor building, image feature-based vSLAM generally has better triangulation performance than Visual SLAM (vSLAM) using disparity information with a lot of noise, such as a stereo camera, and therefore feature-point-based vSLAM generally produces more accurate self-localization results. For this reason, it is best to perform dense measurement of the environment by overlaying stereo information on the estimation results of vSLAM.
[0008] To accurately obtain stereo disparity information of the environment in a dark place, it is necessary to illuminate the environment with patterned lighting. However, since patterned lighting is generally a source of error in feature-based vSLAM and can cause the system to lose its position, it is necessary to ensure that this pattern does not appear in the images used for Visual SLAM.
[0009] In view of the above, an object of the present invention is to provide a mobile body, a mobile body control system, and a mobile body control method that can achieve both dense environmental measurement and highly accurate self-position estimation. [Means for solving the problem]
[0010] In view of the above, the present invention defines a mobile body as "a mobile body including an image sensor capable of acquiring stereo information, an illuminator that provides illumination light by structured lighting and illumination light by normal lighting, an actuator, and a mobile body control system that controls the actuator using image information from the image sensor, wherein the mobile body control system includes an imaging control unit that controls the position and attitude of the illuminator and the illumination light, a self-position estimation unit that estimates a self-position based on feature points determined from image information from the image sensor, a stereo disparity calculation unit that determines stereo disparity information from image information from the image sensor, and an environmental map creation / update unit that creates an environmental map based on the feature points and the stereo information, wherein the mobile body determines feature points using image information from the image sensor obtained by the illuminator providing illumination light by normal lighting, determines stereo disparity information using image information from the image sensor obtained by the illuminator providing illumination light by structured lighting, and creates the environmental map using a combination of the feature points and the stereo disparity information."
[0011] Furthermore, the present invention provides a mobile body control system that is mounted on a mobile body, acquires image information from an image sensor capable of acquiring stereo information, and controls an illuminator that applies illumination light by structured lighting and illumination light by normal lighting, and an actuator, the mobile body control system comprising: an imaging control unit that controls the position and attitude of the illuminator and the illumination light; a self-position estimation unit that estimates a self-position based on feature points determined from image information from the image sensor; a stereo disparity calculation unit that determines stereo disparity information from image information from the image sensor; and an environmental map creation / update unit that creates an environmental map based on the feature points and the stereo information, the mobile body control system being characterized in that the illuminator determines feature points using image information from the image sensor obtained by applying illumination light by normal lighting, determines stereo disparity information using image information from the image sensor obtained by applying illumination light by structured lighting, and creates the environmental map using a combination of the feature points and the stereo disparity information.
[0012] Furthermore, the present invention provides a mobile body control method that is realized using a computer mounted on the mobile body, acquires image information from an image sensor capable of acquiring stereo information, and controls an illuminator that applies illumination light using structured illumination and illumination light using normal illumination, and an actuator, wherein the mobile body control method controls the position, orientation, and illumination light of the illuminator, estimates a self-position based on feature points determined from image information from the image sensor, determines stereo disparity information from the image information from the image sensor, creates an environmental map based on the feature points and the stereo information, determines feature points using image information from the image sensor obtained by the illuminator applying illumination light using normal illumination, determines stereo disparity information using image information from the image sensor obtained by the illuminator applying illumination light using structured illumination, and creates an environmental map using a combination of the feature points and the stereo disparity information. [Effects of the Invention]
[0013] According to the present invention, it is possible to achieve both dense environment measurement and highly accurate self-location estimation. [Brief explanation of the drawings]
[0014] [Figure 1] A diagram showing an example of a survey / work site. [Figure 2A] FIG. 1 is a diagram showing an example of the configuration of a wheeled robot. [Figure 2B] FIG. 1 is a diagram showing an example of the configuration of an autonomous mobile robot. [Figure 3] 10A and 10B are diagrams showing examples of time-series images taken when the camera viewpoint is moved using a normal camera. [Figure 4] 1 is a diagram showing an example of the configuration of a mobile object control system according to a first embodiment of the present invention. [Figure 5] FIG. 2 is a diagram showing the data processing content in the mobile object control system. [Figure 6A] 10A and 10B are diagrams illustrating a time division method as a method for combining normal illumination and structured illumination. [Figure 6B] 10A and 10B are diagrams illustrating a time division method as a method for combining normal illumination and structured illumination. [Figure 6C]10A and 10B are diagrams illustrating a time division method as a method for combining normal illumination and structured illumination. [Figure 6D] A diagram showing normal lighting when moving and structured lighting when stationary. [Figure 7A] 10A and 10B are diagrams illustrating a space division method as a combined method of normal illumination and structured illumination. [Figure 7B] 10A and 10B are diagrams illustrating a space division method as a combined method of normal illumination and structured illumination. [Figure 8A] FIG. 10 is a diagram showing an example of a means for realizing structured illumination when capturing images with a stereo camera. [Figure 8B] FIG. 10 is a diagram showing an example of a means for realizing structured illumination when capturing images with a stereo camera. [Figure 9] FIG. 10 is a flowchart showing the overall processing of the mobile object control system when a time division method is adopted. [Figure 10] 10 is a flow diagram of lighting switching when the time division method is adopted. [Figure 11] FIG. 10 is a flowchart showing the overall processing when a space division method is adopted. [Figure 12A] FIG. 10 is a diagram showing a specific example of the structure when switching between normal illumination and structured illumination in a monocular stereo camera with a rotation mechanism. [Figure 12B] FIG. 10 is a diagram showing a specific example of the structure when switching between normal illumination and structured illumination in a multiple illumination switching type monocular stereo camera. [Figure 13A] 10A and 10B are diagrams showing example images captured by a monocular stereo camera with a rotating mechanism. [Figure 13B] 10A and 10B are diagrams showing example images captured by a plurality of illumination-switchable monocular stereo cameras. [Figure 14] FIG. 2 is a diagram showing an example of the structure of data stored in a data storage unit M1. [Figure 15] FIG. 2 is a diagram showing an example of the data structure stored in an environmental map storage unit M2. [Figure 16] FIG. 2 is a diagram showing a detailed configuration example of a self-position estimation unit C2 according to the first embodiment of the present invention. [Figure 17] 10A and 10B are diagrams showing the processing relationship of the self-position estimation unit C2 in the case of time division and space division. [Figure 18]FIG. 2 is a diagram schematically showing the processing contents of a movement assumption unit C22 and a feature point tracking unit C23 in a self-position estimation unit C2. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0016] Before describing the embodiments of the present invention, the prerequisites for the present invention will be described with reference to FIGS. 1 to 3. FIG.
[0017] Figure 1 shows an example of an investigation or work site. This site is a dark, highly radioactive environment where structure 1, rubble 2, and equipment 3 are located.
[0018] A mobile robot (hereinafter referred to as mobile unit 10) such as the one shown in Fig. 2A and Fig. 2b is entered into a survey / work site to perform environmental surveys and various tasks. Note that Fig. 2A uses wheeled travel 15, while Fig. 2b uses autonomous walking 16 as a means of locomotion, but various types of locomotion are applicable. In either case, mobile unit 10 is equipped with at least a sensor 11 that functions as a camera, lighting 12, a mobile unit control system 13, an actuator 14 for travel control (and various tasks), and a communication unit 17. Note that the power source may be built-in or may be externally powered via a cable.
[0019] 3 is a diagram showing an example of a time series of images taken when the camera viewpoint P is moved using a normal camera in a non-high radiation environment. The diagram shows frame image F0-frame image F5 when the camera viewpoint is P1-P5, as indicated by the arrows in FIG. 1. In contrast, when compared with an example of a time series of images taken when the camera viewpoint P is moved using a low-resolution camera, such as a radiation-resistant camera that can operate in a high-radiation environment, the images taken with the low-resolution camera have a lower resolution than the images taken with a normal camera and contain noise-like information, making it difficult to perform feature-point-based self-location estimation processing in the mobile object control system 13. [Example]
[0020] The first embodiment will be described with reference to Fig. 4 to Fig. 13. Based on the above premise, in the first embodiment of the present invention, a mobile object control system 13 is configured as shown in Fig. 4. The mobile object control system 13, which is the core of the mobile object 10, is configured by a computer having a processing unit C and a memory unit M.
[0021] In terms of the processing functions, the calculation unit C can be said to include a feature extraction unit C1, a self-position estimation unit C2, a key frame detection unit C3, a key frame processing unit C4, an environmental map creation / update unit C5, a mobile object control unit C6, a function switching unit C7, a stereo disparity calculation unit C8, and a shooting control unit C9.
[0022] The memory unit M also includes a data storage unit M1 that stores data as intermediate products (such as feature values D2, movement amounts D3, selected key frames D5, and stereo disparity information D10, which will be described in FIG. 5), an environmental map storage unit M2 that stores the environmental map created by the environmental map creation / update unit C5, and a feature dictionary (feature database) M3 that is referenced when determining features.
[0023] The mobile body 10 is equipped with a sensor 11, a communication unit 17, an actuator 14, and a lighting device 12, and the mobile body control system 13, which forms the core of the mobile body 10, obtains various information and data from the sensor 11 and the communication unit 17 and controls the lighting device 12 and the actuator 14.
[0024] 5 is a diagram showing the data processing content in the mobile object control system 13. In the following figures, parallelograms represent data D, and rectangles represent processing functions in the calculation unit C.
[0025] The data processing content in the mobile object control system 13 shown in Fig. 5 can be roughly divided into a lighting control unit CA and an actuator control unit CB. The present invention is characterized particularly by the lighting control unit CA, and the processing in the actuator control unit CB can be realized by appropriate processing based on information obtained as a result of lighting control, so an example is shown here. Note that the lighting control unit CA will be described in Example 1, and the actuator control unit CB will be described in Example 2.
[0026] Before explaining Figure 5, we will explain the basic concept of lighting control in this invention. In this invention, normal lighting and structured pattern lighting (hereinafter simply referred to as structured lighting) are used in combination. Here, normal lighting refers to irradiating the object with illumination light generated by illuminator 12 mounted on mobile body 10 as is, and the object is then imaged by a camera (stereo camera) serving as sensor 11. Structured lighting refers to irradiating the object with illumination light generated by illuminator 12 mounted on mobile body 10 via a pattern structure, and the object is then imaged by a camera (stereo camera) serving as sensor 11.
[0027] 6A, 6B, and 6C are diagrams illustrating a time-division method as a method for combining normal illumination and structured illumination. In FIG. 6A, normal illumination and structured illumination are alternately performed, and a captured frame (captured image D1) is captured in the mobile object control system 13 each time illumination is performed. In FIG. 6B, the allocation of normal illumination and structured illumination is changed (in the illustrated example, the allocation of normal illumination is large), and illumination is alternately performed, and a captured frame (captured image D1) is captured in the mobile object control system 13 each time illumination is performed. In FIG. 6C, lighting is performed in coordination with the control of the mobile object 10, and images are captured with normal illumination during movement and then with structured illumination after a temporary stop. FIG. 6D is a diagram showing normal illumination during movement and structured illumination when stopped. These lighting methods can be said to be time-division combination methods for combining normal illumination and structured illumination.
[0028] Figures 7A and 7B are diagrams explaining the space division method as a method of combining normal lighting and structured lighting. Figure 7A shows an example of an image captured inside a reactor building, showing rubble and the wall in the background. In the space division method, lighting control is performed so that normal lighting is applied to the rubble areas where many features can be obtained, and structured lighting is applied to the wall areas where few features are obtained (featureless areas) for the space inside the reactor building. Note that the image in Figure 7A is a stereo image, so there is actually another image, but it is not shown here.
[0029] To achieve this, it is necessary to separate the normal illumination area from the structured illumination area. For example, in an environment map such as that shown in Fig. 7B formed in the environment map storage unit M2 (described later), an area with few feature points (featureless area) indicated by dots is extracted on the XY coordinate system, and this is recognized as the structured illumination area. The illumination directed toward this area is designated as structured illumination. In addition, other areas where many feature points are detected are recognized as normal illumination areas, and the illumination directed toward this area is designated as normal illumination.
[0030] 8A and 8B are diagrams showing an example of a means for realizing structured illumination when capturing images with a stereo camera (sensor 11). In Fig. 8A, a pattern structure 12a placed in front of an illuminator 12 has a rotatable rotation mechanism structure, thereby realizing structured illumination when the pattern structure 12a is placed vertically and realizing normal illumination when the pattern structure 12a is placed horizontally. In Fig. 8B, a plurality of illuminators 12 are installed, and pattern structures 12a are placed in front of some of the illuminators 12 while not placed in front of the others using a multiple-unit lighting switching mechanism, thereby switching between structured illumination and normal illumination by switching which illuminator 12 is turned on and irradiating it.
[0031] An example of a means for realizing the combined use of normal illumination and structured illumination is shown in the illumination control unit CA in FIG.
[0032] In the lighting control unit CA in Fig. 5, first, the mobile object control system 13 acquires data D1 (hereinafter simply referred to as frame image) of frame image F in time series, as exemplified in Fig. 3, from the sensor 11. For this reason, the sensor 11 has a camera function, but it is more desirable that it be a three-dimensional camera such as a stereo type camera in order to calculate position and distance, which will be described later.
[0033] The data D1 of this frame image F is mainly processed by the actuator control unit CB, the details of which will be explained in the second embodiment, but the illumination control unit CA controls the illumination when this image is taken.
[0034] In the illumination control unit CA, the function switching unit C7 first generates illumination switching instruction information D8. The illumination switching instruction information D8 is information that instructs that either structured illumination or normal illumination should be continuously performed when switching is not performed, and is information that instructs switching between structured illumination and normal illumination and determines the timing of switching when switching is performed.
[0035] Upon receiving the illumination switching instruction information D8, the imaging control unit C instructs the sensor 11 (stereo camera) on the timing of imaging (not shown), and also provides the illuminator 12 with illumination control information D9.
[0036] The lighting control information D9 at this time includes information on the timing of normal lighting or structured lighting as exemplified in Figures 6A, 6B, and 6C, and by instructing the sensor 11 (stereo camera) on the timing of shooting in synchronization with this, information D1 on the shooting frame is obtained.
[0037] Alternatively, the lighting control information D9 is information on the illumination area of normal lighting or structured lighting when the space is divided as exemplified in FIG. 7A. In this case, the function switching unit C7 refers to an environmental map, such as that shown in FIG. 7B, formed in the environmental map storage unit M2 of FIG. 5, extracts an area with few feature points (featureless area) indicated by dots on the XY coordinate system, and obtains lighting switching instruction information D8 that sets the lighting directed at this area as structured lighting and sets the lighting for other areas where many feature points are detected as normal lighting.
[0038] The stereo disparity calculation unit C9 calculates stereo disparity using input images from the sensor 11 (stereo camera) acquired under structured lighting or normal lighting conditions to obtain stereo disparity information D10. The stereo disparity information D10 corresponds to the distance between the shooting position and the position of the object. The stereo disparity information D10, along with other intermediate product information, is stored in the data storage unit M1 and is used for subsequent position estimation, etc.
[0039] 9 is a flow diagram showing the overall processing of the mobile object control system 13 when the time division method is adopted. According to this diagram, a sensor image D1 is acquired in processing step S11, and a function switching unit D8 generates illumination switching instruction information D8 in processing step S12.
[0040] The processing in processing step S13 is premised on the time division method exemplified in Figures 6A, 6B, and 6C, and when the lighting switching instruction information D8 indicates normal lighting (yes in processing step S13), the following processing is performed within the actuator control unit CB: extracting features (processing step S14), performing self-position estimation processing (processing step S15), performing environmental map creation / update processing (processing step S16), and performing mobile object control output (processing step S17).
[0041] On the other hand, when it is determined in processing step S13 that the lighting switching instruction information D8 indicates structured lighting (No in processing step S13), the stereo disparity calculation unit C9 generates stereo disparity information D10 (processing step S18), performs three-dimensional point cloud processing (processing step S19), and then performs environmental map creation and update processing (processing step S16) as processing within the actuator control unit CB, and performs each process of mobile object control output (processing step S17).
[0042] 10 is a flowchart showing the illumination switching process when the time-division method is employed. According to this diagram, in processing step S21, the function switching unit C7 determines the timing for switching the illumination. Furthermore, in processing step S22, the function switching unit C7 determines whether to issue an illumination switching command. In processing step S23, if the illumination switching command information D8 indicates normal illumination (yes in processing step S23), the imaging control unit C8 executes normal illumination in processing step S24. If the illumination switching command information D8 does not indicate normal illumination (No in processing step S23), the imaging control unit C8 executes structured illumination.
[0043] Fig. 11 is a flow diagram showing the overall processing when the space division method is adopted. According to this diagram, a sensor image D1 is acquired in processing step S31, and structured illumination is identified by function switching unit C7 in processing step S32. At this time, the region estimation of Fig. 7B stored in data storage unit M1 is performed, and the normal illumination region and structured illumination region in the image are identified separately.
[0044] In processing step S33, it is assumed that the spatial division method illustrated in FIG. 7A is used, and when the screen area indicates a normal illumination area (yes in processing step S33), the actuator control unit CB subsequently extracts features (processing step S34), performs self-position estimation processing (processing step S35), performs environmental map creation and update processing (processing step S36), and performs mobile object control output processing (processing step S37).
[0045] On the other hand, if the determination in processing step S33 indicates a structured illumination area (No in processing step S13), the stereo disparity calculation unit C9 generates stereo disparity information D10 (processing step S38), performs three-dimensional point cloud processing (processing step S39), and then performs environmental map creation and update processing (processing step S36) as processing within the actuator control unit CB, and performs each process of mobile object control output (processing step S37).
[0046] The stereo camera in the present invention is preferably realized using a monocular stereo camera. The monocular stereo camera is described in detail in JP 2021-12075, and its general structure is as follows: "A first mirror has a first reflective surface that is a curved surface that is convex in a first direction, has a first vertex, and has a first fan-shaped shape; and a second mirror has a second reflective surface that is convex in a second direction opposite to the first direction, has a second vertex opposite to the first vertex, and has a second fan-shaped shape. The imaging optical system forms an image of a first light that is emitted from the subject and reflected by the first reflective surface, and then reflected by the second reflective surface, and a second light that is emitted from the subject and reflected by the second reflective surface, and receives the light on an image sensor. The interior angle between the first fan-shaped shape and the second fan-shaped shape is 180° or more. The image sensor is positioned so that its center position is shifted from the optical axis of the imaging optical system, and so that the short side of the light-receiving surface of the image sensor is approximately parallel to the center line of the image of the first fan-shaped or second fan-shaped shape."
[0047] In Example 1 of the present invention, the configurations shown in Figures 12A and 12B are proposed as specific structures for switching between normal illumination and structured illumination using the monocular stereo camera. Figure 12A is applied to the rotation mechanism type shown in Figure 8A, in which a ring illumination 12c is placed above the monocular stereo camera 12b, and a pattern structure 12a is placed in a portion of the periphery of the monocular stereo camera. Figure 12B is applied to the multiple illumination switching mechanism type shown in Figure 8B, in which a ring illumination 12c is placed above the monocular stereo camera 12b, and a ring-shaped pattern structure 12a is placed below the periphery of the monocular stereo camera.
[0048] 13A and 13B show examples of images captured by the monocular stereo camera with such a configuration. In the case of FIG. 13A, a portion of the 360-degree surrounding image is an image captured using structured illumination, and the remaining area is an image captured using normal illumination. In the case of FIG. 13B, a portion of the 360-degree surrounding image is an image captured using structured illumination, and the remaining area is an image captured using normal illumination. When a ring illumination 12c is placed above the monocular stereo camera 12b, the image captured is a normal illumination image of 360 degrees around the stereo camera, as shown in the bottom of FIG. 13B. When a ring-shaped pattern structure 12a is placed below the monocular stereo camera, the image captured is a structured illumination image of 360 degrees around the stereo camera, as shown in the top of FIG. 13B. Suggest a configuration.
[0049] According to the present invention described above, it is possible to achieve both dense environment measurement and highly accurate self-location estimation. [Example]
[0050] Returning to FIG. 5, an example of the configuration of the actuator control unit CB in the mobile object control system 13 will be described.
[0051] In the actuator control unit CB in Fig. 5, first, a feature extraction unit C1 extracts multiple feature amounts from one frame image D1 to obtain feature amount data D2 as a group of feature points. Here, feature amounts are assumed to be feature-point-based SLAM (Simultaneous Localization and Mapping), and are preferably extracted using a feature point extraction method such as SIFT (scale invariant feature transform) or ORB (oriented fast and rotated brief). Note that many feature extraction methods are known, and any appropriate method can be applied to the present invention.
[0052] The feature amount data D2 obtained here is stored in the data storage unit M1. Fig. 14 shows an example of the data structure stored in the data storage unit M1, and the feature amount data D2 stored in Fig. 14 includes an INDEX (D11) which is an ID in the order of data acquisition, an acquisition time D12 displayed in year, month, day, and second (yyyymmddhhmmss), an image (image data) D13, and data of a feature point group A (hereinafter referred to as p) which is a group of feature points acquired from one frame image and is a group of vectors in which each feature point has multidimensional information. A It is composed of D14.
[0053] Although not shown in FIG. 14, the data storage unit M1 also stores stereo disparity information D10 calculated by the stereo disparity calculation unit C9. Although the storage format is not shown, the data basically includes information such as INDEX (D11), which is an ID in the order of data acquisition, acquisition time D12 displayed in year, month, day, and second (yyyymmddhhmmss), and image (image data) D13. In addition, the data of the feature point group A (hereinafter referred to as p A ) includes disparity information D10 obtained from one frame image instead of D14 and D17.
[0054] Returning to Fig. 5, the self-position estimation unit C2 tracks feature points at the same location in the environment for each successive image frame to determine the position of the sensor 11 (the position of the moving object 10) and the distance and angle to the feature points. At this time, by continuously processing multiple feature points per frame, the estimation accuracy can be improved. Note that the estimation method of the self-position estimation unit C2 from multiple feature points will be described in detail later, with a specific processing configuration example shown in Fig. 16.
[0055] Similarly, in the case where the acquired information is the disparity information D10, the self-position estimation unit C2 performs self-position estimation from the disparity information D10.
[0056] Thus, data D3 on the position and orientation of the sensor 11 of the moving object 10 is obtained as a processing result in the self-position estimation unit C2. The data D3 on the position and orientation of the sensor 11 of the moving object 10 is also stored in the data storage unit M1. The data D3 on the position and orientation stored in the data storage unit M1 in Fig. 14 includes data D15 on the position t expressed by the vectors x, y, and z, and data D16 on the orientation R expressed by a 3x3 rotation matrix, both linked to INDEX(D11).
[0057] Here, the processing relationship of the self-location estimation unit C2 during the time division described in FIGS. 6A, 6B, and 6C and during the space division described in FIG. 7A will be described using FIG. 17. First, the upper part of FIG. 17 will be described using the example of time division with alternating illumination described in FIG. 6A. The self-location estimation unit C2 performs position estimation (VSLAM processing) from the feature point group for the normal illumination frame obtained at time t. Next, the self-location estimation unit C2 performs position estimation (parallax information processing) from the disparity information D10 for the structured illumination frame obtained at time t+1. At the next timing, similar to time t, the self-location estimation unit C2 performs position estimation (vSLAM processing) from the feature point group for the normal illumination frame obtained at time t+2. In this way, structured illumination is illuminated only during disparity information processing. In other words, frames are temporally divided into those illuminated with structured illumination and those not illuminated, and the position estimation basis data in the self-location estimation unit C2 is separated and calculated.
[0058] When handling the self-localization results and point cloud superposition results during time division, if switching occurs during movement, it is advisable to assume a motion model such as uniform linear motion or uniformly accelerated linear motion, and interpolate and estimate the self-localization at the time of stereo disparity information acquisition based on the time difference. Furthermore, if the shape of the environmental map for feature-point-based vSLAM can be sufficiently acquired, shape matching with the stereo disparity information can be performed to superimpose a dense point cloud based on the stereo disparity information, or manual position and orientation correction can be performed. When stopped, stereo disparity information can be superimposed based on the position and orientation of the previous self-localization result. The dense point cloud of the superimposed stereo disparity information can be smoothed using point clouds within a certain distance at the same position acquired at different times, which is expected to have effects such as improved accuracy and noise reduction.
[0059] Next, regarding the lower part of Figure 17, to explain the example of spatial division mentioned in Figure 7A, since a normal illumination area and a structured illumination area are included in one frame image, the self-position estimation unit C2 performs position estimation from the feature point group (vSLAM processing) for the former in-frame area, and performs position estimation from the disparity information D10 (disparity information processing) for the latter in-frame area.
[0060] Time division and space division can be used not only individually but also in combination to provide more flexible environmental recognition according to the environment. For example, by creating a sparse map of the surrounding environment using vSLAM processing with normal lighting in a time-division manner, and then switching between normal lighting and structured lighting in a space-division manner in the work area or passage area close to the robot, it is possible to obtain dense stereo disparity information for nearby areas while maintaining high accuracy of self-localization even in situations involving unstable movement or work, and to create an environmental map that is more suitable for movement and work.
[0061] Returning to Figure 5, the position and orientation data D3 from the sensor 11 of the moving object 10 is then processed by a key frame detection unit C3, which selects an image from a plurality of chronologically consecutive images as a candidate key frame D4. The key frame detection unit C3 detects frame images to be used for correcting the environmental map and position and orientation as key frame images, and reserves frames detected under certain conditions as candidate key frames D4. The key frame processing unit C4 selects a frame image from the candidate key frame data D4 to be ultimately used for correcting the environmental map and position and orientation as a selected key frame D5.
[0062] As a method for selecting key frames, it is recommended to select frames such as those shown below as key frames. For example, it is recommended to acquire key frames evenly according to the amount of movement, or to select one frame from a group of frames captured stably and continuously with a large number of environmental feature points as the key frame. It is also recommended to select frames in which the viewpoint has changed significantly and the feature points have moved significantly as key frames. It is also useful to extract only relatively clear images with little noise. Selecting key frames in this way can contribute to highly accurate position and orientation detection and highly accurate map creation. The selected key frame D5 is also stored in the data storage unit M1.
[0063] The selected key frame data D5 stored in the data storage unit M1 in FIG. 14 is the data D17 of the feature point group B in the case of normal lighting. The data D17 of the feature point group B is p B These are feature points obtained from one frame image using a feature point tracking algorithm (feature point tracking unit C23 in FIG. 16, which will be described later).
[0064] The environmental map creation / update unit C5 calculates the distance and angle between the sensor 11 and the group of feature points using the data (feature amount data D2, position and orientation data D3, and selected key frame data D5) stored in the data storage unit M1 in Fig. 14, and creates a new map of the site environment as shown in Fig. 1, or updates the map by adding the newly calculated distance data to an existing map. Note that there are many known methods for creating and updating an environmental map using the feature amount data D2, position and orientation data D3, and selected key frame data D5, and any appropriate method can be applied to the present invention.
[0065] The environmental map created by the environmental map creation / update unit C5 is stored in the environmental map storage unit M2, an example of which is shown in FIG. 15. The data constituting the environmental map storage unit M2 are the acquisition time D22, position D23, orientation D24, and feature point group A (or disparity information D10). However, since the environmental map creation / update unit C5 performs processing using selected key frame data D5, these pieces of data are linked to the key frame number D21 of the selected key frame data D5, and furthermore, this data group linked to the key frame number D21 is assigned a flag D26 that distinguishes between true and false values. The flag D26 distinguishes whether the feature point group of this key frame is used in the environmental map (True) or not (False).
[0066] The data stored in the data storage unit M1 and environmental map storage unit M2 formed as described above includes position and orientation information, but in the data storage unit M1, the position and orientation information D15 and D16 is treated as relative information between the sensor 11 and the target object (1, 2, 3 in FIG. 1), while in the environmental map storage unit M2, the position and orientation information D23 and D24 is treated as absolute information on the coordinate system that indicates the on-site environment, which will be convenient in subsequent application situations.
[0067] By the above series of processes in Figure 5, the self-position of the mobile object 10 is obtained in the data storage unit M1. Furthermore, information on the distance between the mobile object and each position in the field environment (which corresponds to a map) is obtained in the environmental map storage unit M2. The mobile object control unit C6 in Figure 5 autonomously determines a movement route based on the self-position stored in the data storage unit M1 and the distance stored in the environmental map storage unit M2, or determines an actuator control value D7 to move along a movement route instructed from the outside via a cable or the like, and the actuator 14 moves the mobile object 10 in accordance with this.
[0068] 16 is a diagram showing a detailed configuration example of a self-location estimation unit C2 according to an embodiment of the present invention. The self-location estimation unit C2 finally outputs data D3 of the amount of movement and data D6 of the self-location estimation failure based on the feature amount D2 for the data obtained under normal lighting and the on-site environment map information stored in the environment map storage unit M2.
[0069] The self-location estimation unit C2 in Fig. 16 is mainly composed of an initialization determination unit C21, a movement assumption unit C22, a feature point tracking unit C23, a feature amount matching unit C24, and a position estimation calculation unit C25. First, the initialization determination unit C21 clears and initializes the information stored in the environmental map storage unit M2 when the mobile object 10 is first introduced into the field environment. Alternatively, if a current field environment map has been acquired from information from another mobile object 10, it is set to the most recently acquired information. Both of these cases are referred to as initialization.
[0070] The self-location estimation unit C2 of the basic configuration according to the embodiment of the present invention performs the following three-stage processing to robustly continue tracking. Note that Fig. 18 is a diagram schematically showing the processing contents of the movement assumption unit C22 and the feature point tracking unit C23 in the self-location estimation unit C2.
[0071] First, the motion assumption unit C22 performs position and orientation estimation based on a motion assumption. The motion assumption assumes, for example, uniform linear motion and estimates the position of the current frame from the amount of movement (amount of change in position and orientation) in consecutive past frames. For example, the amount of change in position and orientation in the past two frames at time t-1 and time t-2, i.e., the velocity per frame, is calculated, and it is assumed that the motion is at the same velocity from time t-1 to the current frame (time t), and the positions of previously observed feature points are estimated from the positions of the position and orientation based on this assumption. If the feature point in the current frame is present within a predetermined range from the estimated position of the feature point, it is considered to have been estimated correctly.
[0072] The left part of Fig. 18 shows the process of the motion assumption unit C22. Here, three consecutive moving object positions (sensor positions) P t-2 and P t-1 and P t , and the image frame F t-2 and F t-1 and F t and illustrate the positions of feature points in the field environment. Points within image frame F are the positions of multiple feature amounts extracted from one frame, and points outside frame F represent the positions of feature points in the field environment. Note that here, the matrix that changes the relative position and orientation between each moving object position P is set to δ.
[0073] Image Frame F t-2 and F t-1 is the moving object position (sensor position) P t-2 and P t-1 The two frames on the left and center are taken from F t-2 and F t-1 is the past information, and the right frame F t In this case, if the position and orientation of the current frame are estimated from the movement of the past two frames on the left side, and the projected point is detected within a predetermined range, high-speed position estimation is possible.
[0074] That is, in the left side of FIG. 18, for example, the sensor position P t-2 Image frame F when capturing the site position fromt-2 The feature point position in and the sensor position P t-1 Image frame F when capturing the site position from t-1 The feature point position in and the sensor position P t-2 and the sensor position P t-1 When the change in relative position and orientation (transformation matrix δ) between the frames is measured, the sensor position P t-1 The image frame F when the site position is captured from the location Pt which is further transformed by the transformation matrix δ t In this method, it is possible to estimate the position of a feature point within a frame, and when a feature can be observed within a predetermined width from the estimated position, the position can be estimated to the corresponding position for each point. In this case, since the moving body (sensor) does not only move in a straight line, estimation must be performed taking into account changes in direction (changes in posture).
[0075] Based on the above idea, the motion assumption unit C22 estimates the position and orientation of the current frame from the past two frames based on the motion assumption of the feature point detection position. Note that when such estimation is performed, there are cases where the estimated position of the next frame cannot be detected, and a flag D6a indicating whether the frame position can be estimated or not is set. The motion assumption unit C22 further estimates the position and orientation of the next frame F. t The position information of the intra-frame feature observed at position Pt in is estimated.
[0076] The right part of Figure 18 shows a schematic diagram of the processing contents of the feature point tracking unit C23. Here, feature points between the previous and current frame images are tracked based on the similarity of the images between the previous and next frames, and 3D measurement is performed using epipolar constraints based on the correspondence of feature points between the two frames to determine the position and orientation. This processing is generally called visual odometry.
[0077] As a specific circuit configuration for this, the feature point extraction unit C231 extracts the previous frame F t-1 The feature points are extracted from the frame F estimated by the motion assumption unit C22. tFor each of these past and current feature sets, the feature point tracking unit C232 extracts position information of the feature in the frame observed at the position Pt in the current frame F. t Track which of the feature points above correspond to each other. For this tracking, it is recommended to use the Lucas-Kanade optical flow method. Here, for example, all combinations of the old and new feature points are generated, and combinations with low likelihood of directionality are eliminated.
[0078] After feature point tracking is performed, multiple feature point pairs (combinations of feature points that correspond to each other in previous and next frames) are used in the visual odometry estimation unit C233 to estimate the position (distance, direction) and orientation between the sensor 11 and the feature points in the on-site environment using the bipolar constraint equation, the five-point algorithm, and RANSAC. The position and orientation estimation results are calculated as information E = (R, t) of position t and orientation R in the fundamental matrix calculation unit C234. As a result of processing in the feature point tracking unit C23, a flag D6b indicating whether or not estimation was possible is set for information that could not be tracked.
[0079] The information E=(R, t) on the position t and orientation R is stored in the data storage unit M1, and then stored in the environmental map storage unit M2 via the environment / map creation unit C5. As a result, the processing results of the feature point tracking unit C23 are reflected in the processing of the movement assumption unit C22, thereby enabling the position / orientation estimation function to be complemented and integrated.
[0080] In the process of FIG. 18, at least the image frame F t-2 and F t-1 This information is obtained from the environmental map storage unit M2, which stores key frame information. By using the feature information from the environmental map storage unit M2, various estimation processes can be performed based on highly reliable information.
[0081] The feature matching unit matches feature points using a common method known as Bag-of-visual words (BoVW). It finds the correspondence between feature points observed in a frame and those observed in a key frame, and then obtains the correspondence between the feature points in the frame and 3D points based on that. If feature matching fails, it sets a flag D6 indicating a self-localization failure.
[0082] It is advisable to utilize the feature dictionary storage unit M3 in this process. Feature quantities can be selected from appropriate locations within the frame, but in this case, locations that are known to be advantageous for focusing on in order to perform various calculations with high precision, as well as their shapes and characteristics, can be compiled into a database and created as a dictionary, which can be used as reference information when comparing new and old feature quantities, thereby improving accuracy.
[0083] 16, as mentioned above, the tracking results in the feature point tracking unit C23 are reflected in the processing in the movement assumption unit C22, thereby improving accuracy. Using the processing results of these feature point tracking unit C23, movement assumption unit C22, and feature amount matching unit C24, data D10 on the amount of change in one posture between the past and present is found, and data D3 on the amount of movement is provided to the position estimation calculation unit C25.
[0084] In addition, when the self-location estimation unit C2 sets the flag D6a or D6b indicating whether self-location estimation is possible or not, or when it issues the self-location estimation failure D6 in the processing of Figure 16, the self-location re-estimation unit C2R extracts a new past keyframe image from the environmental map storage unit M2, selects a new image D1 based on the new past keyframe image, and re-executes the processing from the feature extraction unit C1 to the self-location estimation unit C2, thereby preventing the subsequent processing from being performed using unstable data and obtaining unstable, unreliable results.
[0085] In performing the lighting switching of the present invention described above, the following points may be further considered. First, the distribution of vSLAM processing is made large in terms of time and space. If the self-position estimation accuracy is high, the mapping accuracy will also increase. Therefore, since the priority of self-position estimation accuracy is high, it is necessary to increase the distribution of the processing of the self-position estimation unit in terms of time and space. That is, it is advisable to increase the ratio of the number of frames without irradiating the pattern light by structured lighting or the ratio of the image area for feature point extraction.
[0086] That is, in the case of time division, when the number of frames for pattern light irradiation is N1 and the number of frames without pattern light irradiation is N2, it is advisable to set N1 < N2. Also, in the case of spatial division, between the image area R1 irradiated with pattern light illumination and the image area R2 irradiated with pattern light illumination, it is advisable to set R1 < R2.
[0087] Second, it is advisable to irradiate the pattern light on an area with few feature points. Extract the feature points from the image and irradiate the pattern light in the direction of the area with few feature points. Also, in the environmental map, it is advisable to irradiate the pattern light on an area with few feature points.
[0088] According to the present invention described above, it is possible to achieve both dense environmental measurement and highly accurate self-position estimation.
[0089] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Also, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations may be made.
[0090] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole in hardware, for example by designing them as integrated circuits, or may be realized in software by having a processor interpret and execute a program that realizes each function.
[0091] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC card, an SD card, or a DVD.
[0092] In addition, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily represent all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0093] 10: Moving object 13: Mobile control system C: Arithmetic section M:Memory part C1: Feature extraction unit C2: Self-position estimation part C3: Keyframe detection section C4: Keyframe processing section C5: Environmental Mapping and Update Department C6: Mobile control unit M1: Data storage M2: Environmental map storage unit M3: Feature dictionary
Claims
1. A mobile body including an image sensor capable of acquiring stereo information, an illuminator that provides illumination light by structured illumination and illumination light by normal illumination, an actuator, and a mobile body control system that controls the actuator using image information from the image sensor, the mobile object control system includes an imaging control unit that controls the position and orientation of the illuminator and the irradiated light, a self-position estimation unit that estimates a self-position based on feature points determined from image information from the image sensor, a stereo disparity calculation unit that determines stereo disparity information from the image information from the image sensor, and an environmental map creation / update unit that creates an environmental map based on the feature points and the stereo information, a moving body that determines the feature points using image information from the image sensor obtained by the illuminator applying illumination light using normal illumination, determines the stereo disparity information using image information from the image sensor obtained by the illuminator applying illumination light using structured illumination, and creates the environmental map using both the feature points and the stereo disparity information.
2. The moving body according to claim 1, The moving body is characterized in that the photography control unit performs time-division illumination control to provide a time period in which structured illumination light is applied and a time period in which normal illumination light is applied.
3. The moving body according to claim 1, The imaging control unit uses the created environmental map to divide the area into a structured illumination area to be illuminated with structured illumination and a normal illumination area to be illuminated with normal illumination, and controls the illumination light from the illuminator to be directed toward each of the divided areas, thereby performing illumination control using a space division method.
4. The moving body according to claim 2 or 3, A moving body characterized in that a proportion of illumination not illuminated by structured illumination is increased in a time-division manner or a space-division manner.
5. The moving body according to claim 1, A moving object characterized in that the image sensor capable of acquiring stereo information is a monocular stereo camera.
6. A mobile object control system is mounted on a mobile object, acquires image information from an image sensor capable of acquiring stereo information, and controls an illuminator that provides illumination light by structured illumination and illumination light by normal illumination, and an actuator, the mobile object control system includes an imaging control unit that controls the position and orientation of the illuminator and the irradiated light, a self-position estimation unit that estimates a self-position based on feature points determined from image information from the image sensor, a stereo disparity calculation unit that determines stereo disparity information from the image information from the image sensor, and an environmental map creation / update unit that creates an environmental map based on the feature points and the stereo information, a mobile object control system characterized in that the feature points are determined using image information from the image sensor obtained by the illuminator applying illumination light using normal illumination, the stereo disparity information is determined using image information from the image sensor obtained by the illuminator applying illumination light using structured illumination, and the environmental map is created using a combination of the feature points and the stereo disparity information.
7. 7. The mobile object control system according to claim 6, A mobile object control system characterized in that the photography control unit performs time-division illumination control to provide a time period in which structured illumination light is applied and a time period in which normal illumination light is applied.
8. 7. The mobile object control system according to claim 6, The imaging control unit uses the created environmental map to divide the area into a structured illumination area to be illuminated with structured illumination and a normal illumination area to be illuminated with normal illumination, and controls the illumination light from the illuminator to be directed toward each of the divided areas, thereby performing illumination control using a space division method.
9. 9. A mobile object control system according to claim 7 or claim 8, A mobile object control system characterized in that the proportion of illumination not illuminated by structured lighting is increased in the time-division illumination distribution or the space-division illumination distribution.
10. 7. The mobile object control system according to claim 6, A mobile object control system characterized in that the image sensor capable of acquiring stereo information is a monocular stereo camera.
11. A mobile object control method, which is realized by using a computer mounted on the mobile object, acquires image information from an image sensor capable of acquiring stereo information, and controls an illuminator that provides illumination light by structured illumination and illumination light by normal illumination, and an actuator, comprising: the mobile object control method includes controlling the position and orientation of the illuminator and the irradiated light, estimating a self-position based on feature points determined from image information from the image sensor, determining stereo disparity information from the image information from the image sensor, and creating an environmental map based on the feature points and the stereo information; a mobile object control method for determining the feature points using image information from the image sensor obtained by the illuminator applying illumination light using normal illumination, determining the stereo disparity information using image information from the image sensor obtained by the illuminator applying illumination light using structured illumination, and creating the environmental map using both the feature points and the stereo disparity information.
12. The mobile object control method according to claim 11, A mobile object control method characterized by performing time-division illumination control to provide a time period in which structured illumination is applied and a time period in which normal illumination is applied.
13. The mobile object control method according to claim 11, A mobile object control method characterized by using the created environmental map to divide the structured illumination area to be illuminated by structured illumination and the normal illumination area to be illuminated by normal illumination, and performing illumination control using a space division method in which the illumination light from the illuminator is directed toward each of the divided areas.
14. The mobile object control method according to claim 12 or 13, A mobile object control method characterized in that a proportion of illumination not illuminated by structured illumination is increased in a time-division illumination distribution method or a space-division illumination distribution method.
Citation Information
Patent Citations
Information processor, information processing method, program, and system
JP2019125345A