Robot work vehicle

The imaging device with neural network analysis and LIDAR enhances human detection in agricultural machinery, ensuring safety by stopping the vehicle and preventing theft, overcoming limitations of conventional systems.

JP2025004418A5Pending Publication Date: 2025-12-26ISEKI & CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023104100
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Conventional technologies struggle to accurately identify humans in agricultural and construction machinery operations, especially when they are moving and not wearing specific work clothes, leading to safety concerns.

Method used

The implementation of an imaging device on the vehicle that utilizes a neural network to analyze human face positions, limb movements, and skin color, combined with LIDAR for three-dimensional shape recognition, and adjusts lighting to enhance image analysis, enabling precise human detection and safety control.

Benefits of technology

This system effectively distinguishes between humans and other objects, preventing accidents by stopping the vehicle or disabling operations when humans are detected within dangerous proximity, and reducing false detections through adaptive lighting and neural network analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To prevent an accident by allowing an unmanned robot work vehicle to accurately distinguish a human in a field.SOLUTION: An imaging device is made to have a function that can measure a distance to an object, as well as a movement and shape of the object, and a function that can distinguish a tint of the object, so as to analyze an image of the object using artificial intelligence to distinguish a human, and allow a work vehicle to perform a safe work to prevent an accident.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention distinguishes between people through image analysis of an imaging device, and controls vehicle travel and work equipment. This relates to a work vehicle that performs the above operations. [Background technology]

[0002] Robotic machinery and remote-controlled operation technology are desired for agricultural and construction machinery. In farm fields, people may approach work vehicles, so safety measures are required. It is important to distinguish between the two. (Patent Document 1) [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2022-123742 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional technology recognizes specific work clothes in specific locations and issues an audible alarm to alert workers. There is technology that has controls that provide arousal.

[0005] However, it is difficult to identify a moving object as a human unless the location and clothing are limited. The present invention aims to provide a work vehicle that can use a neural network to identify humans in a field based on the position of their face, the movement of their limbs, and their color, and then safely control the vehicle. [Means for solving the problem]

[0006] The first aspect of the present invention is achieved by the following technical means.

[0007] The imaging device attached to the vehicle body has a function of recognizing the shape of a moving object, a function of measuring the position of the object, and a function of recognizing the color of the object, The function of recognizing color is capable of recognizing human skin color among the colors of the object, When the height position of the human skin color is at a predetermined height or higher from the ground, the imaged object is compared with pre-registered human shape data. .

[0008] The second invention is solved by the following technical means.

[0009] The imaging device attached to the vehicle body has a function of recognizing the shape of a moving object, a function of measuring the position of the object, and a function of recognizing the color of the object, The function of recognizing color is capable of recognizing human skin color among the colors of the object, a size detection unit that detects the height and width of the object, the position and amount of movement of the limbs equivalent portion, and the position and amount of movement of the head equivalent portion, as means for identifying that the imaged object has a limbs equivalent portion and a head equivalent portion that are part of a human shape, and a color detection unit that detects the height position of the color of the human skin and the color of the limbs equivalent portion, The detected data of the size detector and the detected data of the hue detector are matched to generate a state-changed image. .

[0010] The third aspect of the invention is solved by the following technical means.

[0011] A stably fixed portion and a partially oscillating portion are detected from the human-like shape, and the limb-corresponding portion is the partially oscillating portion. .

[0012] The fourth aspect of the present invention is achieved by the following technical means.

[0013] The imaging device is provided with a lighting device near a color identification device with a color recognition function, and the color tone emitted from the lighting device can be modulated. When the image data detected by the size detection unit and the color tone detection unit do not match, the lighting device has the function of modulating the color tone emitted from the lighting device and changing the masking color determination. .

[0014] The fifth aspect of the invention is achieved by the following technical means.

[0015] The color recognition device and the lighting device are disposed in front of and below the operator seat, and are disposed parallel to the ground or facing obliquely downward. The sixth aspect of the present invention is achieved by the following technical means. If the imaged object is determined to be a human, If the distance between the object and the vehicle is less than a first predetermined distance (α), the work equipment is stopped or the vehicle is stopped from moving, and if the distance between the object and the vehicle is less than a second predetermined distance (β) that is shorter than the first predetermined distance (α), the operating function of the vehicle operating unit equipped with a steering wheel is cut off and operations are only accepted via external communication. [Effects of the Invention]

[0016] field The configuration may be such that a human-like shape is determined in the image.

[0017] field The neural network can be used to determine whether the human-like shape in the image is a human or not.

[0018] Size By taking measures when the images detected by the detection unit and the color detection unit do not match, false detection can be prevented.

[0019] human This enables safety control when the distance between the vehicle and the work vehicle is within a dangerous range, and also enables response to accidents such as theft.

[0020] Imaging device This makes it possible to detect the placement position of the object with higher accuracy. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is an overall perspective view of a work vehicle according to an embodiment of the present invention; [Figure 2] FIG. 1 is a left side view of a work vehicle according to an embodiment of the present invention. [Figure 3] 1 is a perspective view of the vicinity of an imaging device of a work vehicle according to an embodiment of the present invention; [Figure 4] FIG. 1 is a top view of the vicinity of an imaging device of a work vehicle according to an embodiment of the present invention; [Figure 5] Skin color matching diagram of a human-shaped figure in the conversion processing of a captured image of the present invention [Figure 6] Anatomical diagram of a human-shaped body in the conversion process of a captured image of the present invention [Figure 7] A linear transformation diagram of a human-shaped figure in the transformation processing of a captured image of the present invention. [Figure 8] Amplitude diagram of a human-shaped arm in the conversion process of a captured image of the present invention [Figure 9] Foot amplitude diagram of a human-shaped object in the conversion process of a captured image of the present invention [Figure 10] A matching diagram of a human-shaped face in the conversion processing of a captured image of the present invention. [Figure 11] A diagram showing the relationship between the color of the face of the human-shaped body and the ground height of the human-shaped body of the present invention. [Figure 12] Relationship between amplitude and ground height of the human-shaped object of the present invention [Figure 13] A diagram showing the relationship between the movement of the main body and limbs of the humanoid shape of the present invention. [Figure 14] FIG. 10 is a diagram showing the relationship between the reflectance of the light amount of the lighting device of the present invention and the height of a human figure. [Figure 15]A block diagram of human judgment and work process change by shape change analysis using neural networks in an embodiment of the present invention. [Figure 16] 10 is a diagram showing another embodiment of the present invention in which an imaging device is disposed on a work machine when the work machine is located below the machine body. [Figure 17] FIG. 10 is a diagram showing another embodiment of the present invention in which the imaging device is disposed on the front frame when the work machine is located below the machine body. [Figure 18] FIG. 10 is a diagram showing another embodiment of the present invention in which an imaging device is disposed in the center of the cabin. [Figure 19] FIG. 10 is a diagram showing another embodiment of the present invention in which imaging devices are arranged at the four corners of the cabin. [Figure 20] FIG. 10 is a diagram showing another embodiment of the present invention in which an imaging device is arranged on a safety frame. [Figure 21] FIG. 10 is a diagram showing another embodiment of the present invention in which an imaging device is disposed on a top link. [Figure 22] A diagram showing an arrangement of a millimeter wave radar and an imaging device in another embodiment of the present invention. [Figure 23] FIG. 10 is a diagram showing another embodiment of the present invention in which an imaging device is disposed below a spare seedling table. [Figure 24] FIG. 10 is a diagram showing an arrangement of an imaging device for checking the working state of a work machine in another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] The present invention will be described below with reference to the embodiments shown in the drawings.

[0023] The work vehicle and related part diagrams shown in Figs. 1 to 15 show an example of this embodiment.

[0024] The background of the present invention will be explained.

[0025] To address the labor shortage in agriculture, unmanned agricultural machinery, known as robotic work vehicles, are being used. These robotic work vehicles are operated under the supervision of a human in a remote location. The main reason for the human supervision is to prevent accidents such as a person approaching the work vehicle and being caught in the vehicle during operation.

[0026] To address this issue, a system is needed in which the robotic work vehicle itself can identify humans and perform safe control. This invention utilizes an imaging device to accurately analyze people and enables the robotic work vehicle itself to perform safe control, thereby realizing an unmanned work vehicle.

[0027] The construction of the work vehicle of the present invention will be described.

[0028] The present invention relates to the configuration and control of an imaging device generally provided on vehicles that travel and work in autonomous agricultural machinery, such as tractors that are fitted with work implements for tilling fields, rice transplanters that plant rice and vegetable seedlings, vegetable transplanters, combine harvesters that harvest rice and vegetables, and dedicated vegetable harvesters.

[0029] The construction of the work vehicle used in this explanation will be explained using a tractor 100 in Figure 1. As an agricultural work vehicle, the tractor 100 is a robotic vehicle that pulls a work implement and is equipped with a satellite positioning device 111 and an inertial positioning device 112. Below this device that confirms the vehicle's position, an imaging device called a LIDAR (Light Detection and Ranging) 120 is installed. This is a type of remote sensing technology that uses light and is capable of measuring scattered light in response to pulsed laser irradiation to analyze the shape and properties of an object. It also makes it possible to measure the distance to an object, from long distances to short distances, and the moving speed of an object.

[0030] The imaging device is provided at the center of the front part of the vehicle body. A left light 140 and a right light 150 are provided at a position behind the vehicle body relative to the position of the CCD camera 130. As shown in Figures 2, 3, and 4, the left light 140 and the right light 150 are positioned above the CCD camera 130 and behind it in the direction of travel of the vehicle. They are located a distance behind, dimensionally corresponding to 142. Therefore, even if the CCD camera 130 has a wide-angle imaging range 131, the illumination range of the left light 140 is 141 and the illumination range of the right light 150 is 151, making it possible to sufficiently illuminate the imaging range 131.

[0031] In FIG. 4, the illumination range 151 is configured symmetrically such that the angle of spread is smaller than the line connecting the position of the end 132 of the CCD camera 130 and the right end 152 of the right lighting 150. This prevents the illumination ranges 141 and 151 from directly illuminating the lens position of the CCD camera 130, preventing halation and enabling the color of the imaged object to be accurately captured.

[0032] Furthermore, the CCD camera 130, which is an imaging device, is disposed below the left lighting 140 and the right lighting 150. This positional relationship also makes it possible to prevent halation.

[0033] It is well known that the positions of the CCD camera and lighting can be moved in the forward / backward, left / right, and up / down directions in this way, but it is important to consider the design of the mobile device when arranging it, and by locating the lighting unit in a recessed position, functionality and design are both achieved.

[0034] The satellite positioning device 111 is disposed above the imaging device 130 and the lighting devices, left light 140 and right light 150, all of which are provided forward of the operation seat 101, or in the form of an operation box, in front of the operation box. The satellite positioning device 111 may be disposed further forward of the operation seat 101 and operation box, but is disposed above or below the operation seat 101 in a position that is not in the field of view of the operation seat 101, and is disposed parallel to the ground or diagonally downward. In the embodiment shown in FIG. 2, the satellite positioning device 111 is located in front of the cabin, which is the operation box.

[0035] A method for analyzing an imaged object using a combination of two types of imaging devices, a CCD camera 130 and a LIDAR 120, will be described.

[0036] 5 shows an image of a person (humanoid figure), and the captured image 200 is an image taken by the CCD camera 130 showing the color and shape. Since the figure is not shown in color, it is shown using diagonal lines. Note that the relationship between the shape of the diagonal lines and the color is not specifically shown, but the difference in the diagonal lines means that the color has changed.

[0037] In this way, the work vehicle has basic data for recognizing humans registered in the data storage unit S15-20 in Figure 15. The basic data includes a large number of registered shape images of human figures, including image data of people walking in a field, standing upright, and standing with their legs spread naturally, and also includes data on the external shape, limb position, head, amplitude range and color of each part. In particular, skin color recognition of the head position is the basis for the human figure search of the present invention. The human figure may be controlled by roughly matching human figures that are similar to the skin color recognition of the head position.

[0038] In human recognition, it is possible to make a highly accurate judgment if the image is set to specific imaging conditions and is easy to analyze, but the present invention is to identify humans in unexpected situations, so first the skin color at the head position of the human-like shape and the human-like shape are matched, and then the search target is focused on, and an image of that target is taken and a detailed analysis is performed; this detailed analysis will now be explained.

[0039] CCD camera 130 may be a monocular camera, but as explained above, its relationship with the lighting device is important for accurate color recognition. In this image, captured image 201 detects a hue corresponding to skin color, but the reflectivity is medium to slightly high, and the object surface has small irregularities, preventing a significant increase in reflectivity. In contrast, captured images 202, 203, and 206 are of colors other than skin color, and compared to captured image 201, they show areas with higher and lower reflectivity. Captured images 204 and 205 may recognize skin color or may detect other colors.

[0040] Captured image 210 is an image of the same object captured by LIDAR 120. The image is point cloud data. In this figure, large circles are displayed for ease of understanding, but in actual images, the points are in the form of tiny dots, and the point cloud shape can also be configured as a smooth shape.

[0041] Shape recognition on a two-dimensional screen is almost the same type of planar recognition as that of the CCD camera 130, but because it is possible to calculate the distance between the object and the LIDAR 120, it is possible to estimate the three-dimensional shape of the captured image 210. In addition, by detecting the movement of the object through point cloud movement detection, it is possible to determine the detailed three-dimensional shape of the captured image 210.

[0042] Therefore, captured image 220 can be synthesized by combining captured image 200 and captured image 210. A feature of the synthesized image is that it allows for coloring of the three-dimensional image. By analyzing the movement of parts, it is possible to identify the parts of an object that move instantly and the parts that remain stably fixed. Furthermore, by analyzing the link-like movement, it is possible to identify the parts that serve as the fulcrum of the link. By analyzing the difference in the speed of this movement and the link configuration, it is possible to distinguish between the main body and the branches and leaves of the object. For example, in the case of a human, it is possible to distinguish between the torso, limbs, and head. In captured image 220, only skin color is extracted and color-recognized in the point cloud data. Captured image 221 is the skin-color recognized part. In this way, captured image 201 taken by the CCD camera is synthesized with captured image 211 by LIDAR to produce captured image 221. The image analysis trend is the same for each part, from captured image 206 ⇒ 216 ⇒ 226 and captured image 205 ⇒ 215 ⇒ 225. This image analysis method is a method for analyzing image data, and corresponds to the first state change, which is the control section of S15-21 in FIG.

[0043] Regarding the recognition of human-like shapes, priority is given to detecting the skin-colored position from the field in the images captured by the LIDAR 120 and CCD camera 130, and if the skin-colored position is at a specified height from the ground in the field and the corresponding object is a moving object, it is analyzed as a human-like shape, and if multiple human-like shapes are recognized, a method is used to analyze each individual.

[0044] Next, referring to Figure 6, the flow of analyzing object movement from image movement will be explained. The left diagram in Figure 6 is the same as the right diagram in Figure 5, and is captured image 220. The case will be explained where captured image 220 moves and becomes captured image 220A. The object's total height 230 changes to total height 230A, but there is no difference in dimensions. Similarly, there is no difference in dimensions between the height of the skin-colored portion 231 and 231A. However, there is a difference in dimensions between 232 and 232A and between 233 and 233A. This difference can be analyzed as a change, either longer or shorter. However, because there is no change in the object's total height, dimensions 232 and 233 are the branches and leaves of the object, which would correspond to the limbs of a person, and this can be analyzed from the change in the first state change.

[0045] By narrowing down the LIDAR image, it is also possible to change the point cloud data into linear data. Captured image 220B in Fig. 7 is an image obtained by linearly converting captured image 220. Similarly, captured image 220A is an image obtained by linearly converting captured image 220C. By converting both images into linear data, the main body part and the peripheral parts can be more clearly distinguished.

[0046] Dimension 232B changes to dimension 232C when the mounting angle changes from dimension 234B to dimension 234C. Dimension 233B changes to dimension 233C when the mounting angle changes from dimension 235B to dimension 235C. However, since dimensions 230B and 230C, which are the overall height, do not change, it can be inferred that dimensions 232B and 233B are parts of branches and leaves that move, and since there is little change or movement in the skin-colored parts between dimensions 231B and 231C, it can be inferred that captured image 220 is a human.

[0047] Figure 8 verifies whether it corresponds to a human arm. In captured image 220D, dimension 232D moves very quickly. It moves smoothly and quickly between 8S1 and 8S5. In contrast, movements of 8S7 to 8S12 are also observed. This movement is more frequent and often faster than the movements of dimension 232E, which is assumed to be a foot in Figure 9, 9S1 to 9S4.

[0048] This technology for judging people using an image capture device can be used to judge the balance and movement of the eyes, nose, mouth, ears, and other parts of the human face, as well as the color of the face, but it is only used to capture images in an easy-to-understand manner so that people can recognize the person in the image. However, when it comes to recognition for the purpose of preventing theft or accidents, there is no way to capture images in an easy-to-understand manner.

[0049] However, the work vehicle used in this invention is for agricultural use and is intended for identification in fields such as rice paddies. Therefore, human movement is often limited. When humans move in fields, they walk on two legs, and it is unlikely that they move while sitting or crawling on all fours. Inside homes, people move in a variety of ways, such as sitting, lying down, or using chairs or desks, resulting in a wide range of changing conditions, making it difficult to identify people. A difficult aspect of identifying people in fields is the wearing of hats, gloves, boots, etc. Furthermore, wearing a towel around the neck or a mask can also make it difficult to identify the head, but walking in fields is often distinctive.

[0050] In the field, people do not sit on the soil, nor do they walk while crouching. Maintaining balance is particularly difficult on soft fields, so people tend to walk with their arms and legs outstretched. This makes it easy to convert the images shown in Figures 5 to 7, and the arm movements shown in Figure 8 and the foot movements shown in Figure 9 are often captured at the intended measurement positions, allowing for a small amount of computation for image analysis and matching. Thus, in the field where the present invention is based, human movement is limited, and analysis from captured images is possible.

[0051] In this way, when the imaging device is equipped with a LIDAR device and a camera that recognizes color, and the position that can be determined to be skin color from the ground is at a certain height, it is clear that it is easy to match human shapes.

[0052] Figures 10, 11, 12, and 13 illustrate control for preventing misidentification of non-human animals. As mentioned above, in agricultural fields, people often wear hats, gloves, boots, towels around their necks, or masks, making it difficult to identify the head. However, in the case of humans, due to movement, it is possible to recognize skin-colored areas by capturing an image from a certain direction. In Figure 10, the face area is calculated by calculating the ratio of the position of dimension 231F to the total height 230F and the ratio of dimension 241 itself. Furthermore, because dimension 241 also moves, the allowable range is also set as dimension 242, which is also determined by the ratio. In this way, regardless of the actual dimensions of the total height 230F, the accuracy of identifying a human can be confirmed by adjusting the ratio.

[0053] Figure 11 shows the subtle differences in skin color depending on the RGB color and overall height. The X axis represents the red hue of RGB, and the Y axis represents the overall height of 230F. Although this differs depending on weather and field conditions, under the same conditions, if the overall height of 230F is small (S11-1), or if the person is short, reds may appear slightly stronger due to the reflected light from the field and the image from the imaging device being slightly downward. Conversely, if the person is tall (S11-2), blues appear slightly stronger, and these color corrections also allow for accurate analysis of the face and its position.

[0054] FIG. 12 shows the relationship between object height and the amplitude of object movement by superimposing consecutively captured object images at object heights 230, 230A, 230B, 230C, 230D, 230E, and 230F. The X-axis represents object height, and the Y-axis represents the object's amplitude. In this embodiment, the image amplitude exhibits large amplitudes at low position 12S-1 and intermediate position 12S-2. Positions 12S-1 and 12S-2 correspond to the position of a person's feet and hands, respectively.

[0055] Furthermore, 12S-3 corresponds to the human torso, and 12S-4 corresponds to the human head, both of which have low amplitude. By synthesizing the outline of an object in this way, the amplitude of specific parts is large, and if their positions match, it is possible to estimate that it may be human. In this embodiment, the solid line indicates the registered reference amplitude value, and the measured value is displayed as a two-dot chain line. The dashed line indicates a large quadrupedal animal, and it can be seen that the trends are different. Note that analysis matching is performed using mathematical means, and if the matching rate is above a predetermined level, it is determined to be a match.

[0056] As mentioned above, such human judgments are action judgments based on specific actions taken by humans in the field. In fields where the soil is soft, such as after plowing, plowing, or immediately after rice planting, bipedal walking often involves stretching out the arms and legs, which increases the accuracy of image judgment. However, on ordinary asphalt roads, the footing is stable, and people often run with bent arms or carry luggage, making it difficult to determine whether they are humans. Control that primarily recognizes people as moving objects is more immediate. However, recognizing people in the field has different requirements. In emergency responses other than work-related entrapment accidents, the prevention of theft and vandalism requires not only accurate recognition of moving objects but also accurate recognition of people.

[0057] For human judgment, it is necessary to increase the degree of external shape matching. This is the analysis of the amount of movement of the limbs mentioned above. The degree of amplitude can be measured as the amount of movement in a simple external image, or by measuring the movement of a specific color. It is also possible to filter out all parts other than the moving parts, then use a convolutional neural network to match specific parts and calculate the degree of overlap. The human shapes that serve as the basis for comparison are registered as cloud data, and can be handled using ICT technology on mobile devices, reducing the load on the work vehicle's CPU.

[0058] Both methods extract the center and head positions of the humanoid shape from the image, calculate the relative and absolute movement amounts between the limbs, and perform image control to improve the recognition of the humanoid shape and thereby increase the accuracy of human identification. This is explained with reference to FIG. 13. Continuously capturing images of the same humanoid shape, over a predetermined period of time, humanoid shape 220G moves to humanoid shape 220H. When the movement amount is measured at the center of the humanoid shape in these images, dimension 251 is detected. Furthermore, a movement amount of dimension 252 is detected for the part estimated to be the humanoid's head. In either case, the movement amount estimated to be the main body of the humanoid shape is small. However, compared to this movement amount, the movement amounts estimated to be the left and right arms are detected as dimensions 253 and 254. Furthermore, the movement amounts estimated to be the left and right feet are dimensions 256 and 257. The movement amount of dimension 257 for the stepping out on the moving side of the human body is particularly large. Thus, dimensions 251 and 252 are considered to be the movement amount of the main body of the humanoid shape and are absolute movement amounts. In comparison, dimensions 253, 254, 256, and 257 are relative movement amounts, which are movement distances including the amount of movement of the main body. In this case, dimensions 253, 254, 256, and 257 are larger than dimensions 251 and 252, which are the amount of movement of the main body, and the parts of dimensions 253, 254, 256, and 257 are presumed to be human limbs. Furthermore, because dimension 251 of the main body center is much smaller than dimension 252, dimension 252 is presumed to be the head, and it is determined that there is a high possibility that humanoid shapes 220G and 220H are human.

[0059] Based on this technology, the fourth invention performs safety control for both the target object and the work vehicle. If it is suspected to be a human, it is extremely dangerous to approach the work vehicle. If the distance to the work vehicle falls below a predetermined distance α, the work equipment will be stopped if it is attached, and then the vehicle will be stopped from traveling. In addition, information on the emergency stop is immediately uploaded to the cloud, allowing users to share the information.

[0060] If the human-like figure approaches within a predetermined distance β even after the vehicle has stopped, this is considered a dangerous act and may be a mischief or theft attempt. In this case, manual operation of the vehicle is blocked. Simultaneous operation is blocked, and operation is only accepted via external communication, i.e., remote operation is the only option.

[0061] If the operating seat is a cabin type, it is also possible to provide mechanical locks such as door locking, steering lock, and fixing of operating levers.

[0062] The values ​​of α and β can be changed depending on the vehicle type and also on the working speed. α is about 4m and β is about 2m, and can be set arbitrarily, but it is recommended to set upper and lower limits, and even if the safety control can be turned off, there is an automatic return function after a certain time has passed.

[0063] In addition to color and movement amount (amplitude), the aforementioned amount of light reflection is also used as a criterion for identifying a humanoid figure. The amount of reflection is indicated by the relative sensitivity of the CCD camera 130, and is calculated by analyzing specific wavelengths of near-infrared light. The left lighting 140 and right lighting 150 are composed of LED lights, and the color is modulated by changing the frequency or wavelength, thereby correcting and controlling the relative sensitivity of the CCD camera 130. When working in the field, the desired light intensity or wavelength may not be obtained depending on the working time or weather. Therefore, the left lighting 140 and right lighting 150 control their LED lights to correspond to the light intensity and wavelength used as the reference for the CCD camera 130.

[0064] As mentioned above, captured image 201 detects a color equivalent to skin color, but the reflectivity is medium to slightly high, and the object surface has small irregularities, preventing a significant increase in reflectivity. In contrast, captured images 202, 203, and 206 are of colors other than skin color, and compared to captured image 201, they show areas with higher and lower reflectivity. The higher reflectivity areas include resin or enamel boots, raincoats, hats, etc. Conversely, the lower reflectivity areas are clothing made from materials such as cotton. Clothing made from synthetic fibers has a higher reflectivity than cotton, but is still low enough to be distinguishable compared to the reflectivity of skin.

[0065] The relative sensitivity of the CCD camera 130 corresponds to the reflectance at a specific wavelength. Therefore, it is possible to identify human skin by extracting the median value of the relative sensitivity. Identifying this median value of relative sensitivity is necessary for each wavelength, and it can be registered in advance in cloud data and identified in the work vehicle by performing a comparison operation.

[0066] Figure 14 shows the relationship between the reflectance of the light intensity of the lighting device and the height of the human figure. The X-axis represents the reflectance of the light intensity of the lighting device, which is similar to the relative sensitivity of the CCD camera 130. The Y-axis represents the height of the human figure. S14-1 is the shoe area. In the case of boots, the reflectance is extremely high. S14-2 is the lower body area, but the reflectance is low due to clothing. S14-3 corresponds to the area from the upper body to the face. The face area has an intermediate reflectance, but can exhibit high reflectance if the clothing is open at the chest or short sleeves. S14-4 is the hair area. In the case of long, black hair, the reflectance is extremely low. Straw hats and other items have rough surfaces, so they are prone to exhibiting a similar tendency. S14-5 is the top edge of the contour, which may be detected as high depending on the angle of sunlight.

[0067] While it is technically possible to identify humans using the reflectance of objects in this way, unlike the prior art in JP 2022-123742, it is not limited to identifying people wearing the same clothing on the premises of a factory or the like. Therefore, in a farm field, if a person is wearing sunglasses or a black mask, has long black hair, and is not wearing a hat, the reflectance will be reduced, which may result in false recognition. Therefore, if a person is recognized as possibly being human, issuing an audio warning is also a necessary safety control measure. For example, if a voice message saying "Danger! Do not approach the vehicle. It may suddenly accelerate," is played, it is expected that a human will respond to the voice.

[0068] In another embodiment, a radiation thermometer may be added to detect the temperature of an object, allowing estimation of the human condition.

[0069] Furthermore, as mentioned above, by modulating the color tone by changing the frequency or wavelength of the lighting device, it is possible to detect even when language cannot be recognized, and even if a large animal other than a human is recognized, it is possible to perform safety control by adjusting the lighting device accordingly.

[0070] The control of identifying humans from images captured by the imaging device described above can be achieved using conventional analysis technology, but doing so inside a work vehicle is difficult as it requires calculations and analysis performed solely by the CPU that controls the vehicle.

[0071] In other words, capturing a large number of images and searching and matching them on a large number of cloud servers places a burden on the CPU and takes time, making it impossible to use for urgent analysis.To solve this problem, it is necessary to control the work machine to make human estimation and judgment using only the work machine, and it is also necessary to have a function that can learn the criteria for this judgment.

[0072] Accuracy can be further improved by using a neural network, an artificial intelligence, as a means of identifying humans from image data. This process is explained in Figure 15, which is a block control diagram of human identification and work process changes through shape change analysis using a neural network.

[0073] In the first invention, the function of measuring the shape, movement, and position of an object based on its distribution is the LIDAR 120, which is the first imaging device S15-1, and the function of recognizing color is the CCD camera 130, which is the second imaging device S15-7. By combining the images from both imaging devices, it is possible to determine whether the captured object is moving or fixed, and to detect specific movement from the size, shape, and color of the object.

[0074] While human detection is an important detection method, because humans wear a variety of clothing, human detection can be estimated based on color and shape, but is not conclusive. Therefore, the present invention performs limited control in a specific location, namely, a field. As described above, when a human moves in a field, they walk on two legs, and the positions of their face and limbs are within a limited range. In reality, the work vehicle works in the field, and its position in the field is confirmed and controlled using a satellite positioning device. As described above, the image data from the LIDAR 120 and CCD camera 130 can detect the overall height and skin color position of a human-like shape from the position on the ground in the field. This allows the work vehicle to roughly determine whether it is a human.

[0075] The second invention further analyzes the human shape.

[0076] The first image capture device S15-1 performs analysis in the humanoid shape size detection unit S15-2 by combining images captured by the LIDAR 120, data from the obstacle sensor 23, and data from various vehicle sensors not shown in this block diagram. A first state change step S15-21 is performed to analyze the humanoid shape using a neural network based on the strength of the connections (synapses) between each piece of data. The analysis calculates the overall height S15-3 of the humanoid shape, the overall width S15-4 of the humanoid shape, the position and amplitude S15-5 of the limb-corresponding parts, and the position and amplitude S15-6 of the head-corresponding parts. This state corresponds to the captured image 210 in Figure 5.

[0077] The second image capture device S15-7 is a CCD camera 130 that recognizes color and shape, and performs a humanoid shape color detection process S15-8. The captured image and data from various vehicle sensors (not shown in this block diagram) are combined, and a first state change process S15-21 is performed to analyze the humanoid shape through the strength of the connections (synapses) between the data using a neural network. This analysis process, which involves changing the image state, calculates the skin color position S15-9, the color corresponding to the limbs S15-10, the masking color determination S15-11, and the relative sensitivity detection S15-12. This state corresponds to the captured image 220 in Figure 5. The masking color determination S15-11 is initially set to a standard setting, but the third invention has a function that automatically changes the setting as it learns.

[0078] By combining the state-changed images from the humanoid shape size detection unit S15-2 and the humanoid shape color detection unit S15-8, the combined image is further analyzed in the second state change S15-22. This image analysis process is the image change process shown in Figures 6 and 7. At this stage, it is possible to recognize and infer a humanoid shape, but further image analysis of the limb amplitude is performed. The basic data in the data storage unit S15-20 is compared with the state-changed image to determine whether the person is a human S15-14.

[0079] This image analysis process is an analytical method that uses a neural network to analyze the image by comparing the movement of the humanoid shape, the movement of the limbs of the humanoid shape, and the amount of movement of the object itself with the amount of movement of the limbs, and by matching the registered data with the state change data analyzed from the imaging device, to identify whether it is a human, and this is determined by the image analysis processes in Figures 8, 9, 10, 11, 12, 13, and 14.

[0080] In this way, the analysis is carried out by taking into account the data obtained from the imaging device and various sensor data, repeatedly changing the image state, comparing it with the data storage unit, registering the differences, and using them for the next data analysis.The analysis method used is an artificial intelligence circuit, or so-called neural network analysis method.

[0081] In using the neural network of the present invention, the series of steps is explained using the block diagram in Figure 15, and the analysis method by changing the state of each image is described above in Figures 5 to 14. This neural network makes it possible to identify people from captured images.

[0082] Furthermore, by utilizing the neural network data of the present invention, the aforementioned emergency stop determination S15-15 and theft prevention determination S15-16 are performed. Furthermore, there are cases where this human recognition is incorrect, and a false recognition estimation S15-17 is always performed in the background control. For example, if an image is recognized as a human, but image data captured a predetermined time after the emergency stop does not show any limb movements that would be considered human, and it is determined that another animal has been mistakenly detected, an operation restart plan S15-41 is made and operation is restarted.

[0083] The learning function of this neural network is to automatically determine the function of the work process change and decision unit S15-30. Although manual operation is also possible by receiving instructions from a mobile terminal or remote control device S15-31, in the present invention, the work process plan S15-37, operation resumption plan S15-41, and vehicle safety control S15-45 are performed by changing the weighting of information based on the learning degree using data from the cloud S15-34, the satellite positioning device S15-32, and the above-mentioned status change information S15-14 to S15-17.

[0084] In one example, in the driving route S15-38, a route is automatically created that does not pass through the location where a human is found, and this location is registered in the map control S15-40. From the next time onwards, a route is created that denies the robot from passing through this location, thereby making the user aware of this.

[0085] In addition, the control timing and accuracy of S15-42 to S15-49 are further adjusted using the learning function, as explained above.

[0086] The basic functions of the third invention will now be described.

[0087] There are cases where the image data detected by the humanoid shape size detection unit S15-2 and the humanoid shape color detection unit S15-8 do not match. In the captured image 220 of Fig. 5, the image of the LIDAR 120 of the area corresponding to the face and the skin-colored area of ​​the CCD camera 130 are misaligned and do not match. This state can occur due to differences in the illuminance of external light, or when the reference color for color recognition is misaligned due to reflected light in the field, or when the background colors, such as the color of the soil in the field and the color of the sky, are similar and discrimination becomes unclear, or when the humanoid shape is moving quickly.

[0088] To address this issue, lighting devices are provided in the imaging device near the color recognition device, and the color tones emitted from the lighting devices can be modulated. In this embodiment, the colors of the right lighting device 140 and the left lighting device 150 can be modulated. When the image detection data from the human-shaped size detection unit S15-2 and the human-shaped color detection unit S15-8 cannot match, the system addresses this issue by changing the modulation of the color tones emitted from the lighting devices and the masking color determination S15-11. The masking color determination S15-11 has a learning function, and by registering data in the field map, if specific changes are required depending on the time and location, the system automatically changes the lighting modulation and masking color determination, thereby automatically addressing this issue in future use.

[0089] As described above, the function shown in Figure 15 combines images from the imaging device and allows analysis to be performed by the calculation function in the work vehicle's CPU, enabling immediate processing and decision-making even if a person suddenly appears in front of the work vehicle. With conventional technology, even if a person suddenly appears, the only recognition possible is that of an obstacle, so while an emergency stop is possible, it is not possible to stop the work equipment or take action against theft. Analyzing and matching images via communication with the cloud takes time, making it difficult for conventional control to immediately anticipate a person. With the present invention, by incorporating artificial intelligence functions, identification and decision-making can be performed within the vehicle, enabling emergency response.

[0090] However, this detection is limited to agricultural work and fields. However, since the images were not taken under conditions where the robot wanted to recognize itself, the conditions for analysis are not easy. To address this, the imaging device is equipped with a LIDAR device and a camera that recognizes color, the color emitted from the lighting device is modulated, and the walking state of people in the field is registered as a database for comparison, making it possible to instantly analyze whether they are human, although this is a guess under limited conditions.

[0091] However, even when matching skin color, there are differences in the imaging conditions and individuals, and even creating a database through extensive learning may not be enough to cover everything. Also, although there is a function to change the matching criteria and tolerance range, there is a high possibility of false recognition in situations where the filtering function is turned off.

[0092] Considering such condition analysis, increasing the accuracy of the judgment criteria for a specific element may result in an inability to make judgment analysis or may even result in misidentification. The invention that solves this problem is the content of the present invention, which makes a comprehensive judgment by measuring and analyzing various elements. In other words, rather than detecting a single condition with high accuracy, it is a method that increases accuracy by combining multiple conditions in a cumulative manner.

[0093] It is desirable to have technology that combines a large number of conditions and processes data in a way that resembles human recognition and judgment, and this is largely due to the unique technology of artificial intelligence, as in the present invention, and is made possible by incorporating control that processes images and responds to changes and outputs in the work process in a way that resembles human judgment.

[0094] This neural network's discrimination function also works outside of farm fields, which are outside the conditions of the first invention. However, when it comes to identifying people, people are in various states, and the accuracy of identifying them as humans decreases. However, while it cannot recognize people, it can recognize them as moving objects, so it can be dealt with by using an imaging device to recognize them as moving objects or obstacles and taking evasive action.

[0095] This function detects foreign objects and obstacles and stops the vehicle when it enters a farm road from a field. Because farm machinery is operating in the field, there is a risk of being caught in an accident, so the accuracy of human identification is necessary, but in other cases, accidents can often be avoided by stopping the vehicle.

[0096] Furthermore, while the speed of vehicles travelling on farm roads is high and theft while in motion is difficult, it is possible for someone to get into a work vehicle while driving at low speed in a field, so human recognition is necessary to deal with theft.

[0097] Next, the mounting positions of the imaging device and satellite positioning device of the present invention for different models will be described.

[0098] 16 shows a model in which the working unit is mounted in the center of the vehicle body, on the underside, and a mower or other work implement 282 is located below the vehicle body. In this model, a stay extends from below the power frame in a position that is not affected by the movement of the work implement, and an image capture device 291 is mounted in front of the work implement and above its uppermost position, so that the area in which the work implement will be working can be captured in advance, and in the case of a mower, this can be used to obtain information such as the height and direction of the grass, the mowing direction and height depending on the grass orientation, and power control.

[0099] Fig. 17 shows a different configuration from Fig. 16, in which the imaging device is disposed in the center of the vehicle body in front of the front axle 301. In Fig. 16, the imaging device is disposed at a low position that matches the height of the work equipment, and can capture images parallel to the ground, but changing the height of the work equipment can make it difficult to measure distances accurately. In Fig. 17, however, the imaging device 302 is disposed as low as possible in a location that can be fixed in position, making it possible to capture images parallel to the ground under conditions where distances can be measured accurately.

[0100] Figure 18 shows the case of imaging from a high position. In order to mount the imaging device in a stable position on top, it is best to install it on top of the cabin 303, which is the operation room. Among these, when measuring while rotating 360 degrees like LIDAR 304, it is best to install it in one place in the center of the cabin.

[0101] Furthermore, when taking images continuously, it is difficult to rotate the image capturing device, so by arranging image capturing devices 305 at the four corners of the cabin as shown in FIG. 19, continuous image capturing in all directions is possible.

[0102] Figure 20 shows the safety frame specifications. To ensure safety in the event of a vehicle tipping over, there is a safety frame 306 that protects the operator up to the head, and this frame is equipped with an imaging device 307. Since the safety frame has a folding function, the folding shape must be taken into consideration.

[0103] Figure 21 shows the mounting position of the rear camera on models without a cabin. There is a link to which the work equipment is attached, and by placing the camera on this link, it is possible to take images that match the movement of the work equipment. In Figure 21, the camera 309 is placed on the top link 308.

[0104] FIG. 22 shows an arrangement in which millimeter-wave radar 260 and imaging device 262 are simultaneously installed. Millimeter-wave radar can be used in agricultural machinery not only to measure the distance to an obstacle but also to detect the obstacle's speed with high accuracy. Therefore, it is effective for detecting the movement of obstacles at the edge of a field, cooperative work, and detecting crops ahead. Since a lower, forward position is more advantageous for crop detection, it can also be installed using a member 263, such as a frame for attaching and detaching the front weight. When the millimeter-wave radar is installed at a lower, forward position, imaging device 262 can be installed at a higher position than the millimeter-wave radar to facilitate use in detecting distant positions. To address the diffuse reflection of sunlight, rainwater, wind, and snow, imaging device 262 can be accommodated in eave-forming portion 261, which is formed by extending the cabin ceiling.

[0105] Figure 23 shows the relationship between the satellite positioning device 272 and the imaging device 271 in the transplanter. The imaging device 272 is supported by a stay 273 of the satellite positioning device 272 and is installed at a lower position than the G satellite positioning device 272. Furthermore, if the transplanter has a control mechanism that detects ridges or obstacles and stops the vehicle, the imaging device 271 is installed below the spare seedling tray 272 and behind the front end 273 of the spare seedling tray. This makes it easy to remove seedling frames and allows the spare seedling tray 272 to protect the imaging device 271 in the event of a collision with an obstacle in front.

[0106] 24 is a diagram showing the position of an imaging device 281 for checking the condition of planted seedlings in a rice transplanter. The imaging device 281 is placed above the planting section 282, at a position equivalent to the upper section of the seedling tank. This makes it possible to deal with mud splashed by the planting section 282 and also makes it easy to check the seedling tank. [Explanation of symbols]

[0107] 100 Tractor 120 Imaging device (LIDAR) 130 Imaging device (CCD camera) 140 Left lighting (LED light) 150 Right lighting (LED light) 200 captured images (CCD camera) 210 Images (LIDAR) 220 Captured image (composite) S15-21 First state change (neural network) S15-22 Second state change (neural network)

Claims

1. An imaging device attached to a moving vehicle body has a function of recognizing the shape of a moving object, a function of measuring the position of said object, and a function of recognizing the color of said object, The function of recognizing color is capable of recognizing human skin color among the colors of the object, A work vehicle having a function of comparing the imaged object with pre-registered human shape data when the height position of the human skin color is at or above a predetermined height from the ground.

2. An imaging device attached to a traveling vehicle body has a function of recognizing the shape of a moving object, a function of measuring the position of said object, and a function of recognizing the color of said object, The function of recognizing color is capable of recognizing human skin color among the colors of the object, a size detection unit that detects the height and width of the object, the position and amount of movement of the limbs equivalent portion, and the position and amount of movement of the head equivalent portion, as means for identifying that the imaged object has a limbs equivalent portion and a head equivalent portion that are part of a human shape, and a color detection unit that detects the height position of the color of the human skin and the color of the limbs equivalent portion, A work vehicle characterized in that a state-changed image is generated by matching the detection data of the size detection unit with the detection data of the color detection unit.

3. A work vehicle as described in Claim 2, characterized in that within the human-shaped form, stably fixed parts and parts that partially oscillate are detected, and the limb-equivalent parts are the parts that partially oscillate.

4. A work vehicle as claimed in claim 2, which is equipped with a lighting device in the vicinity of a color identification device within the imaging device that has the function of recognizing colors, and which is capable of modulating the color emitted from the lighting device and changing the determination of the masking color when the image detection data from the size detection unit and the color detection unit cannot be matched.

5. A work vehicle as described in claim 4, wherein the color recognition device and the lighting device are arranged in front of and below the operator's seat, and the color recognition device and the lighting device are arranged parallel to the ground or facing diagonally downward.

6. When the imaged object is determined to be a human, A work vehicle as described in any one of claims 1 to 5, wherein when the distance between the object and the vehicle is equal to or less than a first predetermined distance (α), the work equipment is stopped or the vehicle stops moving, and when the distance between the object and the vehicle is equal to or less than a second predetermined distance (β) that is shorter than the first predetermined distance (α), the operation function of a vehicle operation unit equipped with a steering wheel is cut off and operations are only accepted via external communication.

Citation Information

Patent Citations

  • Recognition state transmitting device

    JP2022123742A