Skeleton detection device, action prediction device, vehicle, and skeleton detection program

The skeletal structure detection device uses multiple models with varying key points to optimize processing load and accuracy by prioritizing important objects, improving detection performance in neural networks.

JP2025134356APending Publication Date: 2025-09-17DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024032207
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Top-down skeleton detection in neural networks faces a challenge in balancing processing load with the required accuracy of skeleton detection, as reducing the number of keypoints for objects reduces accuracy, while ensuring high-precision detection for all objects is computationally intensive.

Method used

A skeletal structure detection device employs multiple skeleton detection models with varying numbers of key points, assigning objects of higher importance to models with more key points based on image data priorities, optimizing the balance between processing load and accuracy.

Benefits of technology

This approach effectively prioritizes high-precision detection for important objects, balancing processing load and accuracy by assigning objects with higher importance to models with more key points, enhancing the overall detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134356000001_ABST
    Figure 2025134356000001_ABST
Patent Text Reader

Abstract

To optimize balance between a processing load and skeleton detection accuracy (the required number of key points).SOLUTION: A skeleton detection device has a plurality of skeleton detection models mutually differing in number of detected key points. The skeleton detection device uses the plurality of skeleton detection models to detect key points of respective objects of an input image (IN) including images of the plurality of objects based upon image data on the input image. The skeleton detection device sets priority levels for the respective objects based upon the image data on the input image, and allocates each object to one of the plurality of skeleton detection models based upon setting results of the priority levels.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a skeletal structure detection device, a behavior prediction device, a vehicle, and a skeletal structure detection program. [Background technology]

[0002] Skeleton detection (skeleton estimation) using neural networks has been widely put to practical use. There are two types of skeleton detection: top-down skeleton detection and bottom-up skeleton detection, and it is known that the former is more likely to achieve better performance than the latter. However, with top-down skeleton detection, there is a concern that the processing load (computation time) will become too large if the number of objects to be detected in the image becomes too large (see Patent Document 1 below). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-42233 Summary of the Invention [Problem to be solved by the invention]

[0004] In top-down skeleton detection, reducing the number of detected keypoints reduces the processing load, but reducing the number of keypoints reduces the accuracy of skeleton detection (corresponding to the detected keypoints becoming coarser). On the other hand, prioritizing accuracy and performing uniformly high-precision skeleton detection (skeleton detection with a large number of keypoints) for all objects in an image is hardly appropriate in terms of processing load. There is hope for the development of technology that strikes a balance between processing load and the required skeleton detection accuracy.

[0005] The present invention aims to propose a technique for optimizing the balance between the processing load and the required accuracy of skeleton detection (the required number of key points). [Means for solving the problem]

[0006] The skeleton detection device of the present invention has multiple skeleton detection models that detect different numbers of key points, and detects key points of each object in an input image using the multiple skeleton detection models based on image data of the input image containing images of multiple objects.A priority is set for each object based on the image data of the input image, and each object is assigned to one of the multiple skeleton detection models based on the priority setting result. [Effects of the Invention]

[0007] Of multiple objects, it is preferable to preferentially assign objects that are considered to be of high importance to the skeleton detection model with the larger number of key points (the skeleton detection model with the higher accuracy). On the other hand, the importance of each object (for example, its importance to ensuring safety when applied to a vehicle) can be estimated based on the image data of the input image. As with the above-mentioned skeleton detection device, a priority is set for each object based on the image data of the input image, and each object is assigned to one of multiple skeleton detection models based on the priority setting results. This makes it possible to preferentially assign objects that are considered to be of high importance to the skeleton detection model with the larger number of key points. As a result, it is possible to properly balance the processing load and the required skeleton detection accuracy (the required number of key points). [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 2 is a diagram illustrating the relationship between a user and other components according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an internal configuration of an in-vehicle system according to an embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing a shooting area of ​​a camera according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating an internal configuration of an in-vehicle device according to an embodiment of the present invention. [Figure 5] FIG. 2 is a diagram illustrating an internal configuration of a vehicle sensor unit according to the embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing an example of an input image according to the embodiment of the present invention. [Figure 7] FIG. 10 is a diagram illustrating how an input image is divided into a plurality of grid regions according to an embodiment of the present invention. [Figure 8] FIG. 2 is a functional block diagram of a controller in the in-vehicle device according to the embodiment of the present invention. [Figure 9] FIG. 10 is a diagram showing how a bounding box is set for each person in an input image according to an embodiment of the present invention. [Figure 10] FIG. 2 is a diagram showing the structure of an input image sequence according to an embodiment of the present invention. [Figure 11] FIG. 10 is a diagram illustrating the relationship between grid areas and priorities according to an embodiment of the present invention. [Figure 12] FIG. 2 is a diagram showing how object sizes are classified into four size classes according to an embodiment of the present invention. [Figure 13] FIG. 2 is a diagram showing how two skeleton detection models are provided in a skeleton detection unit according to an embodiment of the present invention. [Figure 14] FIG. 1 is a diagram showing a state in which a plurality of people exist in an input image according to a first example belonging to an embodiment of the present invention. [Figure 15] FIG. 2 is a diagram showing how the speed state of a vehicle is classified into three types in the first example pertaining to an embodiment of the present invention. [Figure 16] FIG. 10 is a diagram showing a state in which a plurality of people exist in an input image according to a second example belonging to an embodiment of the present invention. [Figure 17] 10 is a flowchart illustrating an operation of a controller in an in-vehicle device according to a fifth example of an embodiment of the present invention. [Figure 18] FIG. 13 is a diagram showing an input image including an image of a maskable object according to a sixth example of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, examples of embodiments of the present invention will be described in detail with reference to the drawings. In each of the drawings, the same parts are designated by the same reference numerals, and duplicate descriptions of the same parts will be omitted as a general rule. In this specification, for the sake of simplicity, symbols or signs referring to information, signals, physical quantities, functional units, circuits, elements, or components may be used, and the names of the information, signals, physical quantities, functional units, circuits, elements, or components corresponding to the symbols or signs may be omitted or abbreviated.

[0010] FIG. 1 shows the relationship between user U1 and other components assumed in an embodiment of the present invention. User U1 is an occupant of vehicle V1. User U1 is the driver of vehicle V1. Hereinafter, when simply referring to a driver, this refers to the driver of vehicle V1 (hence user U1). However, user U1 may also be an occupant other than the driver (i.e., a passenger in vehicle V1). Vehicle V1 is any type of vehicle. Here, vehicle V1 is assumed to be an automobile or the like that runs on a road. An in-vehicle system 1 is mounted on vehicle V1, and each component of in-vehicle system 1 is installed in an appropriate location in vehicle V1.

[0011] A seat ST1 is installed in the cabin of the vehicle V1. A user U1 sits in the seat ST1. Since it is assumed that the user U1 is the driver, the seat ST1 is the driver's seat. Hereinafter, when simply referring to the cabin, unless otherwise specified, this refers to the cabin of the vehicle V1. Furthermore, below, unless otherwise specified, the inside of the vehicle refers to the internal area of ​​the vehicle V1, and the outside of the vehicle refers to the external area of ​​the vehicle V1.

[0012] The direction from the driver's seat of vehicle V1 toward the steering wheel is defined as "forward," and the direction from the steering wheel of vehicle V1 toward the driver's seat is defined as "rearward." The direction perpendicular to the front-to-rear direction and parallel to the road surface on which vehicle V1 is traveling is defined as the left-to-right direction. The direction perpendicular to the front-to-rear direction and perpendicular to the left-to-right direction is defined as the up-to-down direction. User U1 sits in seat ST1 facing forward. The front-to-rear direction, left-to-right direction, and up-to-down direction correspond to the front-to-rear direction, left-to-right direction, and up-to-down direction as seen from user U1. Unless otherwise specified below, vehicle V1 is assumed to be located on a horizontal road surface, and the traveling direction of vehicle V1 is assumed to be forward (however, the steering angle may be other than zero).

[0013] The world coordinate system is defined as the coordinate system of the real space (actual three-dimensional space) in which the vehicle V1 exists. The world coordinate system is a three-dimensional coordinate system with three mutually orthogonal axes: the WX-axis, the WY-axis, and the WZ-axis. The relationship between the WX-axis, the WY-axis, and the WZ-axis and the front-to-rear, left-to-right, and up-to-down directions is defined as follows: The WX-axis is parallel to the left-to-right direction. The WY-axis is parallel to the front-to-rear direction. The WZ-axis is parallel to the up-to-down direction. The direction from rear to front coincides with the direction from the negative side to the positive side of the WY-axis. The direction from bottom to top coincides with the direction from the negative side to the positive side of the WZ-axis. The direction from left to right coincides with the direction from the negative side to the positive side of the WX-axis (see Figure 3). Here, it is assumed that the vehicle V1 is located on a horizontal road surface, so in the real space, the WX-axis and WY-axis are parallel to the horizontal plane (hence the road surface), and the WZ-axis is parallel to the vertical line.

[0014] 2 shows a schematic block diagram of the in-vehicle system 1. The in-vehicle system 1 includes an in-vehicle device 10, a cruise control device 20, an actuator unit 30, a vehicle sensor unit 40, a camera unit 50, and an HMI 60. The components of the in-vehicle system 1 can transmit and receive any signals and information to and from each other through an in-vehicle network formed in the vehicle V1. The in-vehicle network includes, for example, a CAN (Controller Area Network) and an AVCLAN (Audio Visual Communication Local Area Network).

[0015] The in-vehicle device 10 performs tasks such as detecting the skeleton of a person outside the vehicle (details will be described later). The in-vehicle device 10 may be a drive recorder that cooperates with a camera unit 50 to record the situation outside or inside the vehicle. The driving control device 20 controls the driving of the vehicle V1 using an actuator unit 30. The actuator unit 30 has various driving components such as a motor that realizes the driving of the vehicle V1. Specifically, the actuator unit 30 includes an engine and a motor that generate driving force for the vehicle V1, a steering actuator that drives the steering of the vehicle V1, and a brake actuator that drives the brakes of the vehicle V1.

[0016] The vehicle sensor unit 40 has sensors that detect the details of the driving operation of the vehicle V1 by the driver of the vehicle V1 and sensors that detect various states of the vehicle V1. The vehicle sensor unit 40 outputs vehicle sensor information containing the detection results. The driving control device 20 realizes driving control of the vehicle V1 by driving and controlling the actuator unit 30 in accordance with the vehicle sensor information.

[0017] The camera unit 50 consists of one or more unit cameras that capture images of the outside or inside of the vehicle V1. Each unit camera captures images at a predetermined frame rate. Some of the unit cameras provided in the camera unit 50 are exterior cameras. The exterior cameras have a capture area set outside the vehicle V1, and generate exterior camera images by capturing images of the situation within the capture area. The exterior camera images are images captured of the capture area by the exterior camera. Some of the unit cameras provided in the camera unit 50 are interior cameras. The interior cameras have a capture area set inside the vehicle V1 (i.e., the interior of the vehicle V1), and generate interior camera images by capturing images of the situation in the capture area. The interior camera images are images captured of the capture area by the interior camera. Data representing the content of any image is called image data.

[0018] In the following, attention will be focused mainly on camera 51, which is one of the unit cameras provided in camera unit 50. Camera 51 is an exterior camera that captures images of the front side of vehicle V1. Therefore, as shown in FIG. 3, the capture area of ​​camera 51 includes the front area of ​​vehicle V1. In FIG. 3, hatched area SR1 represents a portion of the capture area of ​​camera 51. However, the capture area of ​​camera 51 may also be the rear area, right side area, or left side area of ​​vehicle V1. The capture area of ​​camera 51 may also include all or part of the front area, rear area, right side area, and left side area of ​​vehicle V1. The front area, rear area, right side area, and left side area of ​​vehicle V1 are areas located in the external area of ​​vehicle V1, in front, rear, right side, and left side of vehicle V1, respectively. Image data of the image captured by camera 51 is sent to in-vehicle device 10.

[0019] The HMI 60 is a human machine interface and is provided with a display device 61, a speaker 62, and an operation input unit 63.

[0020] The display device 61 has a display screen such as a liquid crystal display panel, and displays any video (image) under the control of the in-vehicle device 10, the driving control device 20, or a display control device (not shown). The display device 61 is installed in an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can see the display content of the display device 61. Multiple display devices 61 may be installed in the cabin of the vehicle V1. The display device 61 may be a component of a car navigation system installed in the vehicle V1. The car navigation system may be included in the in-vehicle system 1. The display device 61 may be a display device provided in an information terminal (smartphone, etc.) carried by the user U1.

[0021] The speaker 62 outputs any sound (message, warning sound, music, etc.) under the control of the in-vehicle device 10, the driving control device 20, or an audio device (not shown). The speaker 62 is installed at an appropriate location in the cabin of the vehicle V1 so that each occupant of the vehicle V1 can hear the sound output from the speaker 62. Multiple speakers 62 may be installed in the cabin of the vehicle V1. The speaker 62 may be a speaker provided in the information terminal.

[0022] The operation input unit 63 receives arbitrary operations from each occupant of the vehicle V1. The operation input unit 63 can be configured with operation buttons, a touch panel, or the like. A microphone may be provided in the HMI 60, and voice operations using the microphone may be input to the operation input unit 63. The operation input unit 63 may be an operation input unit provided in the information terminal (smartphone, etc.). In addition, a vibration device that applies vibrations to the occupants (particularly the driver) of the vehicle V1 may be provided in the HMI 60.

[0023] 4 shows the internal configuration of the in-vehicle device 10. The in-vehicle device 10 includes a controller 11, a memory 12, a communication unit 13, and a recording medium 14.

[0024] The controller 11 includes, as hardware resources, a processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), etc. The controller 11 may implement any function, operation, or process that should be implemented by the controller 11 by executing a program recorded in the memory 12 or any other recording medium.

[0025] The memory 12 is configured to include a non-volatile memory such as a ROM (Read Only Memory) or a flash memory, and a volatile memory such as a RAM (Random Access Memory). The memory 12 stores various data referenced by the controller 11 as well as various programs to be executed by the controller 11.

[0026] The communication unit 13 is a communication circuit (communication module) that transmits and receives any signal between the in-vehicle device 10 and a counterpart device different from the in-vehicle device 10. The counterpart device for the communication unit 13 includes components other than the in-vehicle device 10 among the components of the in-vehicle system 1 shown in FIG. 2. The communication unit 13 can communicate with the counterpart device via an in-vehicle network formed in the vehicle V1. The counterpart device for the communication unit 13 can include an external device (such as a server device) connected to an external vehicle network. The external vehicle network includes the Internet and an intranet. Note that the controller 11 can transmit and receive any information to and from the counterpart device using the communication unit 13, but the description of the communication unit 13 may be omitted below.

[0027] The recording medium 14 is a nonvolatile recording medium made up of a magnetic disk, a flash memory, or the like, and stores (records) any information in a nonvolatile manner. The controller 11 is capable of recording any information on the recording medium 14 and reading any information recorded on the recording medium 14. The recording medium 14 may be detachable from the in-vehicle device 10. The recording medium 14 may be external to the in-vehicle device 10 and installed within the in-vehicle system 1. If the in-vehicle device 10 has a drive recorder function, the controller 11 can record image data of an image captured by any unit camera in the camera unit 50 on the recording medium 14.

[0028] FIG. 5 shows an internal block diagram of the vehicle sensor unit 40. The vehicle sensor unit 40 includes sensors 41 to 47. The sensors 41, 42, and 43 are an accelerator pedal sensor, a brake pedal sensor, and a steering wheel sensor, respectively. The vehicle V1 is provided with operational components that receive driving operations from the driver, and the operational components include the accelerator pedal, brake pedal, and steering wheel. The sensors 44, 45, 46, and 47 are a vehicle speed sensor, a steering angle sensor, a G sensor, and a GPS sensor, respectively.

[0029] The accelerator pedal sensor 41 detects the operation of the accelerator pedal of the vehicle V1 by the driver of the vehicle V1, and generates and outputs accelerator pedal operation information indicating the operation of the accelerator pedal. The brake pedal sensor 42 detects the operation of the brake pedal of the vehicle V1 by the driver of the vehicle V1, and generates and outputs brake pedal operation information indicating the operation of the brake pedal. The steering wheel sensor 43 detects the operation of the steering wheel of the vehicle V1 by the driver of the vehicle V1, and generates and outputs steering wheel operation information indicating the operation of the steering wheel.

[0030] The vehicle speed sensor 44 detects the speed of the vehicle V1 and generates and outputs vehicle speed information (vehicle speed pulses) representing the detected speed. The steering angle sensor 45 detects the steering angle (steering angle) of the vehicle V1 and generates and outputs steering angle information representing the detected steering angle. The G sensor 46 detects acceleration acting on the vehicle V1 in a predetermined axial direction and generates and outputs the acceleration detection result as acceleration information. The G sensor 46 may detect acceleration in two mutually orthogonal axial directions or may detect acceleration in three mutually orthogonal axial directions. The GPS sensor 47 receives signals from multiple GPS satellites that form the GPS (Global Positioning System) and generates and outputs vehicle position information based on the received signals. The vehicle position information generated by the GPS sensor 47 represents the current location (current position) of the vehicle V1 using longitude and latitude, or represents the current location of the vehicle V1 using longitude, latitude, and altitude.

[0031] The vehicle sensor information generated by the vehicle sensor unit 40 is information corresponding to the traveling state of the vehicle V1 and includes output information from each sensor within the vehicle sensor unit 40. Therefore, the vehicle sensor information includes accelerator pedal operation information, brake pedal operation information, steering wheel operation information, vehicle speed information, steering angle information, acceleration information, and vehicle position information. However, any of this information may not be included in the vehicle sensor information. Each sensor within the vehicle sensor unit 40 periodically updates the information it should generate, and the latest vehicle sensor information is sequentially output from the vehicle sensor unit 40. Sensors other than sensors 41 to 47 (for example, distance measurement sensors, temperature sensors, rain sensors, illuminance sensors, shift lever sensors, and door lock sensors) may also be provided in the vehicle sensor unit 40.

[0032] FIG. 6 shows an example of an input image IN. The input image IN is a two-dimensional still image obtained by one capture by the camera 51. Image data of the input image IN is supplied from the camera 51 to the in-vehicle device 10, and the input image IN is acquired by the controller 11. Below, the image coordinate system in which an arbitrary input image IN is defined is defined as follows. The image coordinate system is a two-dimensional coordinate system having two axes, the X-axis and the Y-axis, which are orthogonal to each other. The X-axis is parallel to the horizontal direction of the input image IN, and the Y-axis is parallel to the vertical direction of the input image IN.

[0033] In the input image IN, the direction from the negative side of the X-axis to the positive side of the X-axis is rightward, and the direction from the positive side of the X-axis to the negative side of the X-axis is leftward. The rightward and leftward directions in the input image IN correspond to the rightward and leftward directions as seen by the driver of the vehicle V1. With respect to an arbitrary target object located within the shooting area of ​​the camera 51 and in contact with the road surface, if the target object moves rightward in real space as seen by the driver of the vehicle V1, the position of the target object in the input image IN will shift rightward. Conversely, if the target object moves leftward in real space as seen by the driver of the vehicle V1, the position of the target object in the input image IN will shift leftward.

[0034] In the input image IN, the direction from the negative side of the Y axis to the positive side of the Y axis corresponds to the direction away from the vehicle V1 (strictly speaking, the direction away from the camera 51). Therefore, if any target object located within the shooting area of ​​the camera 51 and in contact with the road surface ahead of the vehicle V1 moves from the negative side to the positive side of the WY axis, the position of the target object in the input image IN shifts from the negative side of the Y axis to the positive side of the Y axis.

[0035] The hood, which is part of the body of the vehicle V1, is present within the imaging area of ​​the camera 51, and therefore the hood is reflected in the input image IN. In other words, the input image IN includes an image of the hood of the vehicle V1. The area of ​​the entire image area of ​​the input image IN where the image of the hood of the vehicle V1 exists is referred to as the hood area. In the input image IN of Figure 6, the hatched area corresponds to the hood area. Note that depending on the installation conditions of the camera 51, the body of the vehicle V1 may not be reflected in the input image IN.

[0036] In this embodiment, unless otherwise specified, the road surface refers to the ground surface in real space that the vehicle V1 and pedestrians, etc., come into contact with. Note that the road surface and the road may be interpreted interchangeably. A pedestrian, etc., is a person located on or near the road surface. A person located on or near the road surface is a person who is stationary or who moves on foot. A person located on or near the road surface may be a person who moves on a bicycle, an electric cart, an electric kick scooter, etc.

[0037] The controller 11 can perform region division processing on an arbitrary input image IN to divide the entire image region of the input image IN into multiple image regions. Each image region obtained by division is called a grid region. FIG. 7 shows how the input image IN is divided into multiple grid regions. In the region division processing, the entire image region of the input image IN is divided into M regions in the X-axis direction and N regions in the Y-axis direction. Therefore, the input image IN is divided into a total of (M × N) grid regions. M and N each represent any integer greater than or equal to 2. In FIG. 7, it is assumed that "(M, N) = (11, 11)". An arbitrary grid region is represented by the symbol "R[x, y]". x and y represent any integer. However, x and y in the grid region R[x, y] satisfy "1≦x≦M" and "1≦y≦N". The combined region of a total of (M×N) grid regions R[1,1] to R[M, N] constitutes the entire image region of the input image IN.

[0038] Grid areas R[x,y] and R[x+1,y] are two grid areas adjacent to each other along the X-axis, with grid area R[x+1,y] located on the more positive side of the X-axis than grid area R[x,y]. Grid areas R[x,y] and R[x,y+1] are two grid areas adjacent to each other along the Y-axis, with grid area R[x,y+1] located on the more positive side of the Y-axis than grid area R[x,y]. In a total of (M x N) grid areas, the X-axis direction can be considered the row direction, and the Y-axis direction can be considered the column direction. How to use the area division process will be described later.

[0039] The controller 11 performs top-down skeleton detection. Therefore, the controller 11 performs object detection and then performs skeleton detection on the detected object. A functional block diagram of the controller 11 related to object detection and skeleton detection is shown in FIG. 8. The controller 11 includes functional blocks F1 to F6. The controller 11 is a program execution device (computer) capable of executing any program. All or part of the functions of the functional blocks F1 to F6 may be realized by the controller 11 executing a program recorded in the memory 12 or any other recording medium. Input images generated by sequential photographing with the camera 51 are sequentially input to the controller 11. Each of the functional blocks F1 to F6 can perform its own processing by referring to image data of the input image. Furthermore, the functional blocks F1 to F6 may be capable of referring to any information handled by the controller 11. Note that, with respect to any image, the input, output, recording, saving, generation, and acquisition of an image are synonymous with the input, output, recording, saving, generation, and acquisition of image data of the image. Similar expressions are also interpreted in the same manner. Furthermore, any image-based process or operation is specifically a process or operation based on the image data of that image.

[0040] ---Object detection unit F1--- The functional block F1 is an object detection unit. The object detection unit F1 performs object detection processing on each input image IN. The object detection processing is sometimes simply referred to as object detection.

[0041] In the object detection process, the object detection unit F1 detects whether a detection object exists within a detection target area in the input image IN based on the image data of the input image IN. In any input image or detection target area within the input image, the presence of a detection object in the input image or detection target area specifically refers to the presence of an image of the detection object (in other words, image data of the detection object) within the input image or detection target area. The detection target area may be the entire image area of ​​the input image IN, or a partial area of ​​the entire image area of ​​the input image IN. When the presence of a detection object within the detection target area is detected in the object detection process, the position and shape of the detection object in the input image IN are detected, as well as the type of the detection object. The detection objects detected by the object detection unit F1 include at least a person (human). The detection objects may be only a person. The detection objects may also include animals. In this embodiment, animals refer to vertebrates other than a person (human), such as a dog, cat, cow, or pig. Hereinafter, unless otherwise specified, only a person is considered as the detection object.

[0042] The object detection unit F1 generates and outputs object detection information as a result of object detection processing on the input image IN. The object detection unit F1 supplies the object detection information to functional blocks F2 to F5. The object detection information consists of BBOX information and class information. The BBOX information includes position information of the detected object in the input image IN (information specifying the position of the detected object). In detail, the BBOX information represents the position and shape of the detected object in the input image IN. In the object detection processing, a rectangular area in the input image IN where the image of the detected object exists is specified as a bounding box (hereinafter referred to as BBOX). The BBOX information represents the position and shape of the BBOX. The class information represents the type of the detected object in the input image IN. When multiple detected objects are detected from the input image IN by the object detection processing, object detection information is generated and output for each detected object.

[0043] FIG. 9 shows an input image 600 as an example of the input image IN. When the input image 600 is obtained by capturing an image with the camera 51, objects 610 and 620 are located within the capture area of ​​the camera 51, and as a result, the input image 600 includes images of the objects 610 and 620. The object 610 is a person. Therefore, by performing object detection processing on the input image 600, the position and shape of the object 610 in the input image 600 are detected, and the type of the object 610 is detected as a person. The object 620 is also a person. Therefore, by performing object detection processing on the input image 600, the position and shape of the object 620 in the input image 600 are detected, and the type of the object 620 is detected as a person. In FIG. 9, the area 611 within the dashed rectangular frame on the left is a BBOX set for the object 610, and the area 621 within the dashed rectangular frame on the right is a BBOX set for the object 620. In object detection processing of an input image 600, class information indicating that the type of object 610 is a person is included in the object detection information for the object 610, and class information indicating that the type of object 620 is a person is included in the object detection information for the object 620. Note that if the only object to be detected is a person, the class information may be excluded from the object detection information. Hereinafter, the objects 610 and 620 may be referred to as people 610 and 620.

[0044] The BBOX information for the object 610 specifies the coordinates of one of the four corners of the BBOX 611 (for example, the coordinates of the upper left corner), and the width and height of the BBOX 611. Similarly, the BBOX information for the object 620 specifies the coordinates of one of the four corners of the BBOX 621 (for example, the coordinates of the upper left corner), and the width and height of the BBOX 621. For any BBOX, the width of the BBOX is the length of the BBOX in the X-axis direction, and the height of the BBOX is the length of the BBOX in the Y-axis direction.

[0045] The object detection processing method itself is publicly known, and object detectors that perform object detection processing have been put to practical use. The object detector may be an object detection AI configured with a neural network that has undergone machine learning. In this specification, AI is an abbreviation for artificial intelligence. The object detection unit F1 may include a publicly known object detector (not shown), and the object detection unit F1 may realize object detection processing using the publicly known object detector.

[0046] ---Area division part F2--- The functional block F2 is a region division unit. The region division unit F2 executes the above-mentioned region division process for each input image IN. In the region division process, the region division unit F2 generates and outputs region division information for each input image IN. The region division unit F2 supplies the region division information to the functional block F3. The region division information is information that indicates how the input image IN was divided. The region division information specifies the position, size, and shape of each of the grid regions R[1,1] to R[M,N] in the input image IN.

[0047] The region dividing unit F2 can perform region dividing processing on each input image IN based on the object detection information for that input image IN. At this time, the region dividing unit F2 can dynamically change the values ​​of M and N or dynamically change the size of each grid region based on the object detection information (this will be described later).

[0048] However, the region division unit F2 may always perform a fixed region division process regardless of the object detection information. The fixed region division process described here is hereinafter referred to as a fixed region division process. In the fixed region division process, the values ​​of M and N are fixed to predetermined values, and the position, size, and shape of each of the grid regions R[1,1] to R[M,N] are also fixed according to predetermined conditions. Typically, in the fixed region division process, the entire image region of the input image IN is divided into M equal parts in the X-axis direction and into N equal parts in the Y-axis direction. Since the fixed region division process does not depend on the object detection information, it may be performed before or simultaneously with the object detection process.

[0049] ---Priority setting section F3--- The functional block F3 is a priority setting unit. The priority setting unit F3 executes priority setting processing based on the object detection information and the area division information. In the priority setting processing, the priority setting unit F3 sets a priority for each detection target object detected from the input image IN by the object detection processing. The priority setting unit F3 executes the priority setting processing for each input image IN.

[0050] In the following, unless otherwise specified, it is assumed that a total of n people P[1] to P[n] are detected as multiple detection targets from each input image IN by the object detection process. n represents any integer equal to or greater than 2. The value of n for one input image IN may or may not be the same as the value of n for another input image IN.

[0051] In the priority setting process, the priority setting unit F3 assigns ranks to the persons P[1] to P[n], thereby setting individual priorities to the persons P[1] to P[n]. In this case, the priority setting unit F3 can set a higher priority to some of the persons P[1] to P[n] than to others. The priority setting unit F3 can set different priorities to the persons P[1] to P[n], and in this case, sets one of the first to nth priorities to each of the persons P[1] to P[n]. For any integer i, the i-th priority is higher than the (i+1)-th priority. Therefore, the highest priority is the first priority, and the lowest priority is the n-th priority. However, the priority setting unit F3 may set the same priority to two or more of the persons P[1] to P[n].

[0052] Although a detailed example will be described later, the priority setting unit F3 sets a higher priority to a person among the persons P[1] to P[n] for whom more detailed posture detection or behavior prediction is deemed appropriate based on the object detection information and area division information. The priority setting unit F3 generates priority setting information indicating the result of the priority setting process (priority setting details) and supplies it to the function block F4. The priority setting information indicates the priority set for each of the persons P[1] to P[n].

[0053] ---Bone structure detection section F4--- The functional block F4 is a skeleton detection unit. The skeleton detection unit F4 performs skeleton detection processing on the input image IN based on the image data of the input image IN while referring to the object detection information. The skeleton detection unit F4 performs skeleton detection processing for each input image IN. The skeleton detection processing is sometimes simply referred to as skeleton detection. Note that skeleton detection may also be read as skeleton estimation.

[0054] In the skeleton detection process, the skeleton detection unit F4 detects the position of specific parts of the person in the input image IN for each person detected from the input image IN by the object detection process, and generates and outputs the detection results as skeleton detection information. The skeleton detection unit F4 supplies the skeleton detection information to the functional block F5. There are multiple specific parts. The positions of the specific parts are called key points. Therefore, in the skeleton detection process, each key point for each person in the input image IN is detected. The skeleton detection information indicates each key point detected for each person in the input image IN (i.e., the detected position of each specific part for each person in the input image IN). The connection relationship between the key points is called a skeleton. Skeleton information may also be included in the skeleton detection information.

[0055] The skeleton detection unit F4 is equipped with multiple skeleton detection models. The total number of skeleton detection models provided in the skeleton detection unit F4 can be any number equal to or greater than two. Skeleton detection can be performed using each skeleton detection model. For a common input image IN, the execution timing of the skeleton detection process using one skeleton detection model and the execution timing of the skeleton detection process using another skeleton detection model may be simultaneous or may be shifted from each other. In any one skeleton detection model, the total number of key points detected by the skeleton detection process is referred to as the number of key points. Among multiple skeleton detection models, the number of key points in one skeleton detection model is greater than the number of key points in the other skeleton detection models.

[0056] The three skeleton detection models among the multiple skeleton detection models are high-precision, standard-precision, and low-precision skeleton detection models. For the sake of concrete explanation, the numbers of keypoints in the high-precision, standard-precision, and low-precision skeleton detection models are assumed to be 26, 17, and 9, respectively. The multiple specific parts detected by the high-precision skeleton detection model are the head (between the eyebrows), right eye, left eye, right ear, nose, left ear, neck, right shoulder, center shoulder, left shoulder, right elbow, center spine, left elbow, right wrist, left wrist, right hand (palm of the right arm), left hand (palm of the left arm), right hip, center hip, left hip, right knee, left knee, right ankle, left ankle, right foot, and left foot. The multiple specific parts detected by the standard-accuracy skeleton detection model are the head (between the eyebrows), nose, neck, right shoulder, center shoulder, left shoulder, right elbow, center spine, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, and left ankle. The multiple specific parts detected by the low-accuracy skeleton detection model are the head (between the eyebrows), right shoulder, left shoulder, right wrist, left wrist, right hip, left hip, right ankle, and left ankle. However, the specific parts are not limited to the examples given here. Furthermore, depending on the posture of the person in the input image IN, it may be impossible to detect some key points.

[0057] The method of skeleton detection processing itself is publicly known, and skeleton detectors that perform skeleton detection processing have been put to practical use. The skeleton detector may be a skeleton detection AI configured with a neural network that has undergone machine learning. Each skeleton detection model may be configured using a publicly known skeleton detector (not shown).

[0058] The skeleton detection unit F4 assigns one of a plurality of skeleton detection models to each person in the input image IN based on the priority setting information, and performs skeleton detection processing using the assigned skeleton detection model. For example, if a high-precision skeleton detection model is assigned to person P[1] among persons P[1] to P[n] in the input image IN, skeleton detection processing is performed using the high-precision skeleton detection model for person P[1]. As a result, a total of 26 key points are detected for person P[1]. Also, for example, if a low-precision skeleton detection model is assigned to person P[2] among persons P[1] to P[n] in the input image IN, skeleton detection processing is performed using the low-precision skeleton detection model for person P[2]. As a result, a total of 9 key points are detected for person P[2].

[0059] The skeleton detection unit F4 preferentially assigns high-precision skeleton detection models to people who are set to a relatively high priority, rather than to people who are set to a relatively low priority. There is an upper limit to the total number of people to which high-precision skeleton detection models can be assigned for one input image IN. For this reason, people who are set to a relatively low priority are less likely to be assigned high-precision skeleton detection models than people who are set to a relatively high priority, and are more likely to be assigned standard-precision or low-precision skeleton detection models.

[0060] ---Behavior Prediction Section F5--- The functional block F5 is a behavior prediction unit. The behavior prediction unit F5 executes a behavior prediction process that predicts the behavior of each person in the input image IN through posture detection of each person in the input image IN based on the skeleton detection information while referring to the object detection information. Behavior prediction information indicating the prediction result in the behavior prediction process is generated by the behavior prediction unit F5. The behavior prediction unit F5 supplies the behavior prediction information to the functional block F6.

[0061] Figure 10 shows multiple input images IN arranged in time series. i The input image IN obtained by the camera 51 is particularly the input image IN i For any integer i, time t i+1 is time t iThe time is later than the time of m input images IN1 to IN m The set of input images IN SEQ Here, m represents an integer of 2 or more. m Each of these images contains the image of a person P[1] to P[n], and the input images IN1 to IN m By object detection processing for each of the input images IN1 to IN m In the behavior prediction process, the behavior prediction unit F5 generates object detection information for each of the input images IN1 to IN m Based on the object detection information and skeleton detection information generated for each of m The behavior prediction unit F5 predicts the behavior of the persons P[1] to P[n] at a later time and generates behavior prediction information indicating the prediction result. SEQ The system may identify and track each person within the system.

[0062] For example, time t m In the behavior prediction process at time t m It can be predicted whether person P[1] will stop, move away from vehicle V1, or jump out onto the predicted driving path of vehicle V1 after time t m Based on previous vehicle sensor information, time t m The predicted route at time t m The prediction of the travel path of the vehicle V1 is performed by the controller 11.

[0063] Behavior predictors that predict human behavior using skeletal structure detection have been put to practical use. The behavior predictor may be a behavior prediction AI configured with a neural network that has undergone machine learning. The behavior prediction unit F5 may include a known behavior predictor (not shown), and the behavior prediction unit F5 may realize behavior prediction processing using the known behavior predictor.

[0064] ---Safety Support Department F6--- The functional block F6 is a safety support unit. The safety support unit F6 executes safety support processing to support the safe driving of the vehicle V1 based on the behavior prediction information. The safety support processing may include notification to the user U1. The optional notification may be a notification by displaying a video on the display device 61 (a notification that affects the user U1's vision) or a notification by outputting a sound from the speaker 62 (a notification that affects the user U1's hearing), or a combination thereof. If the HMI 60 includes the vibration device, the optional notification may include a notification by generating a vibration from the vibration device (a notification that affects the user U1's tactile sense).

[0065] For example, consider a case where the behavior prediction unit F5 predicts that person P[1] will run out onto the predicted driving path of vehicle V1 or predicts that person P[1] may run out onto the predicted driving path of vehicle V1. In this case, the safety support unit F6 involved in the safety support processing issues a warning notification to notify user U1 of the prediction result. In this case, the safety support unit F6 involved in the safety support processing may cause the driving control device 20 to perform driving control of vehicle V1 (e.g., control to reduce the speed of vehicle V1 to zero) to avoid contact between vehicle V1 and person P[1].

[0066] ---Grid area and priority relationships--- The relationship between grid areas and priority will be described with reference to FIG. 11. FIG. 11 shows how a total of (M×N) grid areas R[1,1] to R[M,N] are set for an input image IN. In the example of FIG. 11, "(M,N)=(11,11)". FIG. 11 shows an example of an input image IN when a vehicle V1 is traveling on a road that stretches in a straight line ahead of the vehicle V1. However, the content of the input image IN varies widely.

[0067] In any input image IN, the image of the hood of vehicle V1 appears in some of the grid areas R[1,1] to R[M,N]. In the input image IN of FIG. 11, the hatched area corresponds to the hood area. The controller 11 sets the grid areas in which the image of the hood of vehicle V1 appears as invalid areas. At least, the grid areas in which only the image of the hood of vehicle V1 appears are set as invalid areas. Grid areas in which the image of the hood of vehicle V1 appears mixed with other images may also be set as invalid areas. In the following, it is assumed that in any input image IN, the image of the hood of vehicle V1 appears only in a total of (M × 2) grid areas R[1,1] to R[M,1] and R[1,2] to R[M,2]. As a result, the grid areas R[1,1] to R[M,1] and R[1,2] to R[M,2] are set as invalid areas, and the other grid areas do not correspond to invalid areas. Depending on the installation conditions of the camera 51, the hood of the vehicle V1 may not appear in the input image IN, and as a result, an invalid area may not be set.

[0068] The object detection unit F1 may set the invalid area. The object detection unit F1 detects whether a detection target object exists within a detection target area in the input image IN, and at this time, excludes the invalid area from the detection target area. Alternatively, the priority setting unit F3 may set the invalid area. In this case, a person may be detected from the invalid area during the object detection process, mainly due to erroneous detection, but the priority setting unit F3 always assigns the lowest priority to a person detected during the object detection process and located in the invalid area. In the following, it is assumed that a person will not be detected from an invalid area.

[0069] The priority setting unit F3 also sets some of the grid areas R[1,1] to R[M,N] as low-priority fixed areas. In this embodiment, a total of (M × 2) grid areas R[1,N] to R[M,N] and R[1,N-1] to R[M,N-1] are set as low-priority fixed areas, and the other grid areas are not considered low-priority fixed areas. Images of the sky or mountains located far enough away from the vehicle V1 appear (or often appear) in the low-priority fixed areas. The priority setting unit F3 always assigns the lowest priority to people detected in the object detection process who are located in the low-priority fixed areas. This is because even if a person is detected in a low-priority fixed area, it is considered that there is little need for skeleton detection and behavior prediction for the person in the low-priority fixed area.

[0070] Within the entire image area of ​​the input image IN, an area that is different from both the invalid area and the low-priority fixed area is called a dynamic priority area. Therefore, the entire image area of ​​the input image IN coincides with the combined area of ​​the invalid area, the low-priority fixed area, and the dynamic priority area. In this embodiment, only each grid area R[x,y] that satisfies "1≦x≦M" and "3≦y≦N-2" is considered to belong to the dynamic priority area.

[0071] The controller 11 may set the invalid region, the low-priority fixed region, and the dynamic priority region based on known information including camera parameters, etc. The known information includes camera installation information of the camera 51. The camera installation information of the camera 51 indicates the installation position of the camera 51 with respect to the vehicle V1 and the mounting angle of the camera 51. The installation position of the camera 51 indicates the height of the camera 51 from the road surface (the bottom end of the vehicle V1) and the distance between the front end of the body of the vehicle V1 and the camera 51. The mounting angle of the camera 51 indicates the depression angle or elevation angle of the camera 51 with respect to the horizontal plane. The mounting angle of the camera 51 may also represent the Euler angles (pitch angle, roll angle, and yaw angle) of the camera 51. The camera parameters are information for specifying the camera coordinate system for the camera 51 together with the camera installation information of the camera 51, such as the focal length and size of the image sensor of the camera 51. The controller 11 may set the invalid region, the low-priority fixed region, and the dynamic priority region using instance segmentation of the input image IN.

[0072] In this embodiment, unless otherwise specified, any person described below is a person detected from the input image IN by the object detection process. The any person detected by the object detection process is referred to as a person of interest, and a method for setting a priority for the person of interest will be described. The priority setting unit F3 identifies a grid area in which the person of interest is located based on the BBOX information of the person of interest while referring to the area division information. The grid area to which the reference position of the BBOX set for the person of interest belongs is identified as the grid area in which the person of interest is located. The reference position of the BBOX is the midpoint of the bottom edge of the BBOX. The BBOX is a rectangular area, and of the four sides of the rectangle that defines the rectangular area, two sides are parallel to the X-axis and the remaining two sides are parallel to the Y-axis. Of the first two sides (i.e., the two sides parallel to the X-axis), the center of the side located on the negative side of the Y-axis (i.e., the vehicle V1 side) is the midpoint of the bottom edge of the BBOX. However, the reference position of the BBOX may be the center position of the BBOX.

[0073] The input image IN in FIG. 11 includes images of people 660, 670, and 680 within the dynamic priority region, and therefore, the people 660, 670, and 680 are detected by the object detection process. BBOXes 661, 671, and 681 are BBOXes set for the people 660, 670, and 680, respectively. Because the reference position of BBOX 661 is within grid region R[4,5], the priority setting unit F3 identifies grid region R[4,5] as the grid region in which the person 660 is located. Similarly, because the reference position of BBOX 671 is within grid region R[5,7], the priority setting unit F3 identifies grid region R[5,7] as the grid region in which the person 670 is located. Similarly, because the reference position of BBOX 681 is within grid region R[8,5], the priority setting unit F3 identifies grid region R[8,5] as the grid region in which the person 680 is located.

[0074] In real space, the person of interest is located in front of the vehicle V1. As the person of interest moves away from the vehicle V1 along the WY axis (see FIG. 1) in real space, the position of the person of interest in the input image IN shifts along the Y axis from the negative side of the Y axis to the positive side of the Y axis. Therefore, the distance in real space between the person located in the grid region R[x,y] and the vehicle V1 increases as the value of y increases. Regarding the persons 660 and 670 in FIG. 11, the distance between the person 660 and the vehicle V1 is shorter than the distance between the person 670 and the vehicle V1 in real space.

[0075] Comparing grid area R[x,y] and grid area R[x+i,y+p], grid area R[x,y] corresponds to the near grid area, and grid area R[x+i,y+p] corresponds to the far grid area. In grid area R[x+i,y+p], i represents any integer satisfying "1≦x+i≦M", and p represents any natural number satisfying "y+p≦N". Of any two grid areas whose positions in the Y-axis direction are different from each other, one corresponds to the near grid area and the other corresponds to the far grid area. In the input image IN, a person located in the near grid area is referred to as a near person (near object), and a person located in the far grid area is referred to as a far person (far object). The near grid area is a grid area in which an image of an object that is relatively closer to vehicle V1 (distance in real space) than the far grid area appears. That is, in real space, the distance between the closer person and the vehicle V1 is shorter than the distance between the farther person and the vehicle V1.

[0076] In the example of input image IN in FIG. 11 , when comparing grid regions R[4,5] and R[5,7], grid region R[4,5] corresponds to a near grid region, and grid region R[5,7] corresponds to a far grid region. Therefore, when comparing people 660 and 670, person 660 corresponds to a near person, and person 670 corresponds to a far person. Similarly, in the example of input image IN in FIG. 11 , when comparing grid regions R[8,5] and R[5,7], grid region R[8,5] corresponds to a near grid region, and grid region R[5,7] corresponds to a far grid region. Therefore, when comparing people 680 and 670, person 680 corresponds to a near person, and person 670 corresponds to a far person. When two people in the input image IN are in the relationship of a close person and a distant person, the priority setting unit F3 can set different priorities between the close person and the distant person depending on the speed of the vehicle V1 (details will be described later).

[0077] On the other hand, an image of an object located to the right of vehicle V1 in real space appears in the input image IN closer to the grid regions R[M,1] to R[M,N] than to the grid regions R[1,1] to R[1,N]. Conversely, an image of an object located to the left of vehicle V1 in real space appears in the input image IN closer to the grid regions R[1,1] to R[1,N] than to the grid regions R[M,1] to R[M,N].

[0078] The grid area in which an image of the right area outside the vehicle V1 appears is called the right grid area, and the grid area in which an image of the left area outside the vehicle V1 appears is called the left grid area. A straight line that passes through the center or center of gravity of the vehicle V1 body and is parallel to the WY axis (the axis extending from the rear to the front of the vehicle V1) is called the vehicle body centerline. The vehicle body centerline is parallel to the road surface and perpendicular to the left-right direction. A grid area in which an image of an object located outside the vehicle V1 and to the right of the vehicle body centerline in real space appears is the right grid area. A grid area in which an image of an object located outside the vehicle V1 and to the left of the vehicle body centerline in real space appears is the left grid area. A grid area in which an image of an object located outside the vehicle V1 and on the vehicle body centerline in real space appears is the center grid area, and does not correspond to either the right grid area or the left grid area.

[0079] Among the grid areas R[1,1] to R[M,N], some correspond to the right grid areas, some to the left grid areas, and the rest to the center grid areas. In the example of FIG. 11 where "(M,N)=(11,11)" is assumed, any grid area R[x,y] that satisfies "1≦x≦5" corresponds to the left grid area, and any grid area R[x,y] that satisfies "7≦x≦11" corresponds to the right grid area. The same applies to the examples described below (e.g., FIGS. 14, 16, and 18) where "(M,N)=(11,11)" is assumed. Each grid area R[x,y] that satisfies "x=6" corresponds to the center grid area. The controller 11 (priority setting unit F3) may regard each grid area located to the left as a left grid area and each grid area located to the right as a right grid area when viewed from a line passing through the center of the input image IN and parallel to the Y axis. The controller 11 (priority setting unit F3) may regard each grid area through which a straight line that passes through the center of the input image IN and is parallel to the Y axis as a central grid area.

[0080] In the input image IN, a person located in the right grid area is referred to as a right person (right object), and a person located in the left grid area is referred to as a left person (left object). In the example of Fig. 11, person 680 corresponds to a right person, and persons 660 and 670 correspond to left people. The priority setting unit F3 can set different priorities between right people and left people depending on the path of the vehicle V1 (depending on whether it turns right or left) (details will be described later).

[0081] The priority setting unit F3 sets a priority for each person based on the position of each person in the input image IN, the size of each person in the input image IN, and the running state of the vehicle V1. The position of each person in the input image IN is specified by which grid area each person is located in.

[0082] The priority setting unit F3 performs a size classification process that classifies the size of each person in the input image IN into a plurality of size classes based on the BBOX information. The total number of size classes can be arbitrary as long as it is 2 or more. Here, for the sake of concretizing the explanation, the total number of size classes is 4, and the plurality of size classes are assumed to be large size, medium size, small size, and extremely small size, which are the first to fourth size classes. The priority setting unit F3 derives the area S BBOX of the BBOX of the target person based on the BBOX information of the target person, and compares the area S BBOX with predetermined threshold values TH1 to TH3. The area S BBOX is represented by the total number of pixels located within the BBOX of the target person in the input image IN. The threshold values TH1 to TH3 satisfy "0 < TH1 < TH2 < TH3". Fig. 12 shows the relationship between the area S BBOX and each size class. The priority setting unit F3 determines which of the first inequality "S BBOX < TH1", the second inequality "TH1 ≤ S BBOX < TH2", the third inequality "TH2 ≤ S BBOX < TH3", and the fourth inequality "TH3 ≤ S BBOX " is satisfied. Then, the priority setting unit F3 classifies the size of the target person as extremely small size when the first inequality holds, small size when the second inequality holds, medium size when the third inequality holds, and large size when the fourth inequality holds.

[0083] Examples of the driving state of the vehicle V1 related to the setting of the priority include the speed state of the vehicle V1 (see the first embodiment described later, etc.). Also, examples of the driving state of the vehicle V1 related to the setting of the priority include the traveling direction state of the vehicle V1 (a state representing whether it goes straight, turns right, or turns left) (see the second embodiment described later, etc.).

[0084] ​​Below, among the multiple embodiments, some specific operation examples, application techniques, modified techniques, etc. relating to each functional block in FIG. 8 will be described. The matters described above in this embodiment are applied to each of the following embodiments unless otherwise specified and unless there is a contradiction. If there are any matters in each embodiment that contradict the matters described above, the description in each embodiment may take precedence. Furthermore, unless there is a contradiction, the matters described in any of the multiple embodiments shown below can also be applied to any other embodiment (i.e., any two or more of the multiple embodiments can be combined).

[0085] 13, unless otherwise specified, it is assumed that the number of skeleton detection models provided in the skeleton detection unit F4 is two, and that skeleton detection unit F4 is provided with skeleton detection models F4_H and F4_L. Skeleton detection model F4_H is the high-precision skeleton detection model described above, and skeleton detection model F4_L is the low-precision skeleton detection model described above. When skeleton detection model F4_H is assigned to a person of interest in input image IN, skeleton detection model F4_H performs skeleton detection for the person of interest based on image data in the BBOX of the person of interest in input image IN. When skeleton detection model F4_L is assigned to a person of interest in input image IN, skeleton detection model F4_L performs skeleton detection for the person of interest based on image data in the BBOX of the person of interest in input image IN.

[0086] <<First Example>> A first embodiment will be described. In the first embodiment, it is assumed that the above-mentioned fixed region division process is performed in the region division unit F2 and that "(M, N) = (11, 11)". In the first embodiment and the second embodiment described later, the upper limit of the number of people that can be assigned to the high-precision skeleton detection model F4_H for one input image IN is the high-precision upper limit number MAX H In the first embodiment and the second embodiment described later, the value "MAX" is used for the sake of concreteness of explanation. H= 5". Furthermore, in the first embodiment, attention is focused on the speed state of the vehicle V1 as the traveling state of the vehicle V1, and a method for setting a priority according to the speed of the vehicle V1 will be described. In the first embodiment, it is assumed that the vehicle V1 is traveling straight ahead (a state in which the steering angle of the vehicle V1 is zero).

[0087] As described above, people P[1] to P[n] are detected from each input image IN by the object detection process. H ", all of the persons P[1] to P[n] can be assigned to the highly accurate skeleton detection model F4_H. However, in the first embodiment, "n>MAX H " Therefore, out of a total of n people P[1] to P[n], five people can be assigned to the high-accuracy skeleton detection model F4_H, while the remaining people are assigned to the low-accuracy skeleton detection model F4_L.

[0088] FIG. 14 shows one input image IN assumed in the first embodiment. A method for setting priorities for the input image IN in FIG. 14 will be described. In the input image IN in FIG. 14, "n=12," that is, the input image IN includes images of people P[1] to P

[12] . In the input image IN in FIG. 14, people P[1] and P[2] are located in grid area R[4,5], people P[3] and P[4] are located in grid area R[8,5], and people P[5] and P[6] are located in grid area R[4,6]. In the input image IN in FIG. 14, people P[7] and P[8] are located in grid area R[8,6], people P[9] and P

[10] are located in grid area R[5,7], and people P

[11] and P

[12] are located in grid area R[7,7].

[0089] The priority setting unit F3 sets a position evaluation value EV POS and size evaluation value EV SZ Through the derivation of the priority evaluation value EV O The priority evaluation value EV O The higher the position evaluation value EV[i] is, the higher the corresponding priority is. POS , size evaluation value EV SZ, priority evaluation value EV O Specifically, the symbols "EV POS [i]” , “EV SZ [i]”, “EV O The priority setting section F3 is represented by the priority evaluation value EV O [i] is derived according to the following formula (1A) or formula (1B). EV O [i]=EV POS [i]×EV SZ [i] (1A) EV O [i]=EV POS [i]+EV SZ [i] (1B)

[0090] The priority setting unit F3 calculates a size evaluation value EV based on the size of the person P[i] in the input image IN. SZ [i] is determined (derived). The size of the person P[i] is classified into large, medium, small, or very small by the size classification process. When the size of the person P[i] is classified into large, medium, small, or very small, the size evaluation value EV SZ [i], respectively, the value VAL1_EV SZ , VAL2_EV SZ , VAL3_EV SZ , VAL4_EV SZ Here, set “VAL1_EV SZ >VAL2_EV SZ >VAL3_EV SZ >VAL4_EV SZ ≧0”. When the high-precision skeleton detection model F4_H is applied to a person whose size is relatively large in the input image IN, it is easy to accurately detect each keypoint because of the person's large size. On the other hand, even if the high-precision skeleton detection model F4_H is applied to a person whose size is relatively small in the input image IN, it is difficult to accurately detect each keypoint because of the person's small size. In addition, the posture, etc. of a person whose size is sufficiently small is likely not to be very important for behavior prediction or safety support. Taking these into consideration, the larger the size, the lower the size evaluation value EV SZThis makes it easier to actually apply high-precision skeleton detection to people whose size merits high-precision skeleton detection.

[0091] The priority setting unit F3 sets a grid priority value for each grid region in the input image IN according to the speed of the vehicle V1. The priority setting unit F3 identifies the speed of the vehicle V1 from the latest vehicle speed information obtained at the time of setting the grid priority value, and sets the grid priority value based on the identified speed. The grid priority value set for the grid region R[x,y] is represented by the symbol "GP_R[x,y]". Each grid priority value has a positive value. However, some grid priority values ​​may be zero.

[0092] The priority setting unit F3 identifies each grid area in which the persons P[1] to P[n] are located based on the BBOX information of the persons P[1] to P[n] while referring to the area division information. The grid area to which the reference position of the BBOX set for the person P[i] belongs is identified as the grid area in which the person P[i] is located.

[0093] The priority setting unit F3 then sets the grid priority value set for the grid area in which the person P[i] is located as the position evaluation value EV POS That is, if person P[i] is located in the grid region R[x,y], the position evaluation value EV POS The grid priority value GP_R[x,y] is set for [i]. Therefore, in the example of FIG. 14, the position evaluation value EV POS [1] and EV POS The grid priority value GP_R[4,5] is set for person P[2]. Similarly, the position evaluation value EV POS [3] and EV POS The grid priority value GP_R[8,5] is set in [4]. The position evaluation value EV POS is also set similarly.

[0094] The priority setting unit F3 sets a higher priority evaluation value EV OTherefore, for example, "EV O [1]>EV O If "EV[2]" holds, then person P[1] will have a higher priority than person P[2], and "EV[2]" will hold. O [1] <EV O If “EV[2]” is true, then person P[2] will have a higher priority than person P[1]. O [1]=EV O [2]”, a common priority is set for persons P[1] and P[2]. The same applies to combinations other than the combination of persons P[1] and P[2].

[0095] When the speed of the vehicle V1 is relatively slow (in a low-speed state described below), the priority setting unit F3 sets the grid priority value GP_R[x,y] higher than the grid priority value GP_R[x,y+p] (where q is an arbitrary natural number). As a result, when the speed of the vehicle V1 is relatively slow, a higher priority is likely to be set for a closer person than for a more distant person. Conversely, when the speed of the vehicle V1 is relatively fast (in a high-speed state described below), the priority setting unit F3 sets the grid priority value GP_R[x,y+p] higher than the grid priority value GP_R[x,y] (where p is an arbitrary natural number). As a result, when the speed of the vehicle V1 is relatively fast, a higher priority is likely to be set for a more distant person than for a closer person.

[0096] An example of a method for setting a grid priority value according to the speed of the vehicle V1 will be described. The priority setting unit F3 classifies the speed of the vehicle V1 into two or more speed states. Referring to FIG. 15, the speed of the vehicle V1 is assumed to be classified into three stages. That is, the priority setting unit F3 classifies the speed of the vehicle V1 into a low speed state, a medium speed state, and a high speed state. The low speed state is when the speed of the vehicle V1 is lower than the threshold speed TH SP1 The high-speed state is the state in which the speed of the vehicle V1 is greater than or equal to the threshold speed TH SP2 The above is the state. In the medium speed state, the speed of the vehicle V1 is greater than or equal to the threshold speed TH SP1 greater than the threshold speed TH SP2 The threshold speed TH SP1 and THSP2 is a predefined speed, satisfying "0 < TH SP1 < TH SP2 ". The state where the speed of the vehicle V1 is zero may also belong to the low-speed state.

[0097] In the low-speed state, the priority setting unit F3 sets the grid priority value for the nearby grid area higher than the grid priority value for the distant grid area. In the high-speed state, the priority setting unit F3 sets the grid priority value for the distant grid area higher than the grid priority value for the nearby grid area. At this time, for example, if the grid area R[x, 5] is regarded as the nearby grid area, then the grid areas R[x, 6] and R[x, 7] correspond to the distant grid areas (where x is an integer from 1 to 11). Also, for example, if the grid area R[x, 6] is regarded as the nearby grid area, then the grid area R[x, 7] corresponds to the distant grid area (where x is an integer from 1 to 11).

[0098] Therefore, in the low-speed state, by paying attention to the rows of "y = 5", "y = 6", and "y = 7", "GP_R[x, 5]> GP_R[x, 6]> GP_R[x, 7]" holds (where x is an integer from 1 to 11). As a result, in the low-speed state, "EV POS [1]= EV POS [2]= EV POS [3]= EV<​​​​​​​​​​​​​​​​​​​​​​​POS [2]=EV POS [3]=EV POS [4] <EV POS [5]=EV POS [6]=EV POS [7]=EV POS [8] <EV POS [9]=EV POS

[10] =EV POS

[11] =EV POS

[12] ” holds true.

[0100] In the medium speed state, the priority setting unit F3 may set the same value to all grid priority values. Alternatively, the method for setting grid priority values ​​in the medium speed state may be the same as the method for setting grid priority values ​​in the low speed state or the method for setting grid priority values ​​in the high speed state. A method intermediate between the method for setting grid priority values ​​in the low speed state and the method for setting grid priority values ​​in the high speed state may be adopted as the method for setting grid priority values ​​in the medium speed state.

[0101] By using the above setting method, in a low-speed state, the position evaluation value EV POS is the position evaluation value EV of a person located in the distant grid area (distant person). POS In the example of FIG. 14, if we look at the rows of "y=5" and "y=7", the position evaluation values ​​EV of the nearby people (P[1] to P[4]) in the low-speed state are POS is the position evaluation value EV of each distant person (P[9]~P

[12] ) POS Therefore, in a low-speed state, the priority evaluation value EV O is the priority evaluation value EV of distant people (P[9]~P

[12] ) O In a low-speed state, the vehicle V1 should pay more attention to avoiding contact with a nearby person than with a distant person. Therefore, in a low-speed state, the position evaluation value EV POS is higher than the distant person.

[0102] At low speeds, the priority evaluation value EV O is the priority evaluation value EV of a distant person O However, in a low speed state (i.e., when the speed of the vehicle V1 is lower than the threshold speed TH SP1 In the first speed range below), if the near person and the far person are classified into the same size class, the priority evaluation value EV O is the priority evaluation value EV of a distant person O As a result, a higher priority is set for a nearby person than for a distant person. In the example of FIG. 14, person P[1] corresponds to a nearby person, and person P[9] corresponds to a distant person. For example, if people P[1] and P[9] are both classified as medium size in a low-speed state, the priority evaluation value EV O [1] is the priority evaluation value EV of person P[9] O As a result, a higher priority is set for person P[1] than for person P[9].

[0103] In addition, in the high-speed state, the position evaluation value EV POS is the position evaluation value EV of the person (nearby person) located in the nearby grid area. POS In the example of FIG. 14, if we look at the rows of “y=5” and “y=7”, the position evaluation values ​​EV POS is the position evaluation value EV of each nearby person (P[1]~P[4]) POS Therefore, in the high-speed state, the priority evaluation value EV O is the priority evaluation value EV of nearby people (P[1]~P[4]) OWhen the vehicle V1 is traveling at high speed, more attention should be paid to avoiding contact with distant people than with nearby people. Whether the vehicle is traveling at high speed or low speed, it is necessary to take action to avoid contact with all people, but it is thought that the speed of the vehicle V1 will be increased only after the safety of nearby people is sufficiently ensured. On the other hand, when the vehicle V1 is traveling at high speed, if the distance (distance in real space) between the vehicle V1 and the person of interest is not large enough, the necessary braking may not be possible in time. For this reason, when the vehicle is traveling at high speed, the position evaluation value EV POS is higher than the nearby person.

[0104] In high-speed situations, depending on the size of the near and far people, the priority evaluation value EV O is the priority evaluation value of nearby people, EV O However, in a high speed state (i.e., when the speed of the vehicle V1 is higher than the threshold speed TH SP2 In the above-mentioned second speed range), if the distant person and the near person are classified into the same size class, the priority evaluation value EV O is the priority evaluation value of nearby people, EV O As a result, a higher priority is set for distant people than for close people. In the example of FIG. 14, person P[9] corresponds to a distant person, and person P[1] corresponds to a close person. For example, if people P[9] and P[1] are both classified as medium size in a high-speed state, the priority evaluation value EV O [9] is the priority evaluation value EV of person P[1] O As a result, a higher priority is set for person P[9] than for person P[1].

[0105] In addition, when two or more people are located in a common grid area, a person who is classified into a relatively large size class among the two or more people is assigned a higher priority evaluation value EV than a person who is classified into a relatively small size class. Ois derived and a high priority is set. This makes it possible to actually apply or easily apply high-precision skeleton detection to people whose size merits the application of high-precision skeleton detection. For example, in FIG. 14, if person P[1] is classified as large size and person P[2] is classified as medium size, "EV O [1]>EV O [2]”, so a higher priority is set for person P[1] than for person P[2].

[0106] Although the method described here is to variably set each grid priority value in three stages according to the speed of the vehicle V1, each grid priority value may also be changed continuously according to the speed of the vehicle V1.

[0107] The skeleton detection unit F4 detects the highest priority of the people P[1] to P[n] by selecting the highest priority of the people P[1] to P[n]. H The high-precision skeleton detection model F4_H is assigned to only a few people, and the low-precision skeleton detection model F4_L is assigned to the remaining people. H =5". For example, in FIG. 14, consider a case where the first to fifth priorities are set for persons P[1] to P[5] and the sixth to twelfth priorities are set for persons P[6] to P

[12] . In this case, the skeleton detection model F4_H is assigned to persons P[1] to P[5], and the skeleton detection model F4_L is assigned to persons P[6] to P

[12] .

[0108] Priority Evaluation Value (EV) O If it is not possible to distinguish between the fifth and sixth priorities based on the above criteria alone, the distinction may be made based on other criteria. For example, in FIG. 14, O [1]>EV O [2]>EV O [3]>EV O [4]>EV O [5]=EV O [6]” and the priority evaluation value EV O [6] is the priority evaluation value EV O [7]~EV O

[12] Consider the case where the priority evaluation value EVO [5] and EV O Based on [6] alone, it is not possible to distinguish between person P[5] and person P[6] to whom the fifth priority should be set. In this case, the priority setting unit F3 may make the distinction based on other indicators. For example, the reference position of the BBOX of person P[5] is compared with the reference position of the BBOX of person P[6], and the fifth priority is set to the person corresponding to the reference position closer to the center of the input image IN.

[0109] High precision upper limit MAX H Separately, low precision upper limit MAX L The upper limit of low precision MAX may also be determined (this is also true for other embodiments including the second embodiment). L represents the upper limit of the number of people that can be assigned to the low-accuracy skeleton detection model F4_L for one input image IN. In this case, "n>MAX H +MAX L ", then the (MAX H +MAX L ) priority, no skeleton detection model is assigned to the person, and skeleton detection is not performed (this also applies to the other embodiments including the second embodiment).

[0110] <<Second Example>> A second embodiment will be described. In the second embodiment, it is assumed that the above-mentioned fixed region division process is performed in the region division unit F2 and that "(M, N) = (11, 11)". In the second embodiment, as described above, the high precision upper limit number MAX H is 5. In the second embodiment, attention is focused on the course state of the vehicle V1 as the traveling state of the vehicle V1, and a method of setting a priority according to the course of the vehicle V1 will be described.

[0111] As described above, people P[1] to P[n] are detected from each input image IN by the object detection process. H ", all of the persons P[1] to P[n] can be assigned to the highly accurate skeleton detection model F4_H. However, in the second embodiment, "n>MAX H" Therefore, out of a total of n people P[1] to P[n], five people can be assigned to the high-accuracy skeleton detection model F4_H, while the remaining people are assigned to the low-accuracy skeleton detection model F4_L.

[0112] FIG. 16 shows an input image IN assumed in the second embodiment. A method for setting priorities for the input image IN of FIG. 16 will be described. In the input image IN of FIG. 16, "n=12," that is, the input image IN includes images of people P[1] to P

[12] . In the input image IN of FIG. 16, people P[1] and P[2] are located in grid area R[8,5], people P[3] and P[4] are located in grid area R[8,6], and people P[5] and P[6] are located in grid area R[7,7]. In the input image IN of FIG. 16, people P[7] and P[8] are located in grid area R[4,5], people P[9] and P

[10] are located in grid area R[4,6], and people P

[11] and P

[12] are located in grid area R[5,7].

[0113] As in the first embodiment, the priority setting unit F3 sets a position evaluation value EV POS and size evaluation value EV SZ Through the derivation of the above, the priority evaluation value EV O The size evaluation value EV is derived. SZ The method of deriving is the same as in the first embodiment.

[0114] However, the priority setting unit F3 according to the second embodiment sets a grid priority value for each grid area in the input image IN according to the path of the vehicle V1. When the movement vector of the vehicle V1 has a forward component, the traveling state of the vehicle V1 can be broadly classified into a straight traveling state (a state in which the steering angle is zero) in which the vehicle V1 travels straight, a rightward traveling state, and a leftward traveling state. The rightward traveling state is a state in which the vehicle V1 travels while turning right (but may stop temporarily). The leftward traveling state is a state in which the vehicle V1 travels while turning left (but may stop temporarily).

[0115] The priority setting unit F3 may determine whether the traveling state of the vehicle V1 belongs to a straight-ahead state, a rightward traveling state, or a leftward traveling state, based on steering angle information (see FIG. 5) that indicates the steering angle of the vehicle V1.

[0116] When the vehicle V1 travels, the steering angle of the vehicle V1 is reflected in the acceleration information of the vehicle V1 (see FIG. 5). That is, the acceleration of the vehicle V1 changes depending on the steering angle of the vehicle V1. The acceleration that reflects the steering angle of the vehicle V1 is the acceleration of the vehicle V1 in the horizontal direction (i.e., the acceleration in the WX-axis direction and the acceleration in the WY-axis direction). Therefore, the priority setting unit F3 may determine whether the traveling state of the vehicle V1 belongs to a straight-ahead state, a rightward traveling state, or a leftward traveling state based on the acceleration information of the vehicle V1 (the acceleration of the vehicle V1 in the WX-axis direction and the WY-axis direction).

[0117] The priority setting unit F3 identifies each grid area in which the persons P[1] to P[n] are located based on the BBOX information of the persons P[1] to P[n] while referring to the area division information. The grid area to which the reference position of the BBOX set for the person P[i] belongs is identified as the grid area in which the person P[i] is located. Then, the priority setting unit F3 sets the grid priority value set for the grid area in which the person P[i] is located as the position evaluation value EV POS [i]. This is the same as in the first embodiment. Also, as in the first embodiment, the priority setting unit F3 determines (derives) the higher priority evaluation value EV O A higher priority is set for a person associated with the same name.

[0118] When the vehicle V1 is in a state of moving rightward, the priority setting unit F3 sets the grid priority value for the right grid area higher than the grid priority value for the left grid area. As a result, in a state of moving rightward, a higher priority is likely to be set for a person located in the right grid area (person on the right) than for a person located in the left grid area (person on the left). Conversely, when the vehicle V1 is in a state of moving leftward, the priority setting unit F3 sets the grid priority value for the left grid area higher than the grid priority value for the right grid area. As a result, in a state of moving leftward, a higher priority is likely to be set for a person located in the left grid area (person on the left) than for a person located in the right grid area (person on the right).

[0119] For the example of FIG. 16, grid regions R[8,5], R[8,6], and R[7,7] correspond to the right grid region, and grid regions R[4,5], R[4,6], and R[5,7] correspond to the left grid region.

[0120] Therefore, in the rightward moving state, the grid priority values ​​GP_R[8,5], GP_R[8,6], and GP_R[7,7] are higher than the grid priority values ​​GP_R[4,5], GP_R[4,6], and GP_R[5,7]. Simply put, in the rightward moving state, the priority setting unit F3 may set a first grid priority value for all grid areas corresponding to the rightward grid areas and a second grid priority value for all grid areas corresponding to the leftward grid areas. Here, the first grid priority value is greater than the second grid priority value. In the rightward moving state, the priority setting unit F3 may set a difference in the grid priority value among all grid areas corresponding to the rightward grid areas, or may set a difference in the grid priority value among all grid areas corresponding to the leftward grid areas, depending on the reference indicator. The reference indicators include the speed of the vehicle V1, the steering angle of the vehicle V1, and the acceleration of the vehicle V1 in the WX-axis direction and the WY-axis direction, and the turning direction of the vehicle V1 can be identified using the reference indicators.

[0121] Conversely, in the leftward moving state, the grid priority values ​​GP_R[4,5], GP_R[4,6], and GP_R[5,7] are higher than the grid priority values ​​GP_R[8,5], GP_R[8,6], and GP_R[7,7]. Simply put, in the leftward moving state, the priority setting unit F3 may set the first grid priority value for all grid areas that correspond to the left grid areas, and the second grid priority value for all grid areas that correspond to the right grid areas. In the leftward moving state, the priority setting unit F3 may set a difference in the grid priority value among all grid areas that correspond to the left grid areas, or may set a difference in the grid priority value among all grid areas that correspond to the right grid areas, in accordance with the reference indicator.

[0122] In a straight-ahead state, the priority setting unit F3 does not set a difference between the grid priority value for the left grid area and the grid priority value for the right grid area. Simply put, in a straight-ahead state, the priority setting unit F3 may set a common grid priority value for all grid areas in the input image IN. In a straight-ahead state, the priority setting unit F3 may set a difference in the grid priority value among all grid areas in the input image IN depending on the speed of the vehicle V1, etc.

[0123] By using the above setting method, in the rightward moving state, the position evaluation value EV POS is the position evaluation value EV of the person located in the left grid area (left person). POS In the example of FIG. 16, the position evaluation values ​​EV[1] to P[6] of the people on the right in the rightward moving state are POS The position evaluation value EV of each of the people P[7] to P

[12] corresponding to the person on the left POS Therefore, in the rightward moving state, the priority evaluation value EV O is the priority evaluation value EV of the person on the left (P[7]~P

[12] ) O In the vehicle V1 traveling to the right, more attention should be paid to avoiding contact with the person on the right than with the person on the left. Therefore, in the vehicle V1 traveling to the right, the position evaluation value EVPOS is higher than the person on the left.

[0124] In the rightward moving state, depending on the size of the person on the right and the person on the left, the priority evaluation value EV O is the priority evaluation value EV of the person on the left O However, if the person on the right and the person on the left are classified into the same size class in the right-moving state, the priority evaluation value EV O is the priority evaluation value EV of the person on the left O As a result, a higher priority is set for the person on the right than for the person on the left. In the example of FIG. 16, person P[1] corresponds to the person on the right, and person P[7] corresponds to the person on the left. For example, if people P[1] and P[7] are both classified as medium size in the rightward moving state, the priority evaluation value EV O [1] is the priority evaluation value EV of person P[7] O As a result, a higher priority is set for person P[1] than for person P[7].

[0125] In addition, in the leftward moving state, the position evaluation value EV POS is the position evaluation value EV of the person located in the right grid area (right person). POS In the example of FIG. 16, the position evaluation values ​​EV[7] to P

[12] corresponding to the left-hand person in the leftward moving state are POS is the position evaluation value EV of P[1] to P[6] corresponding to the person on the right. POS Therefore, in the leftward moving state, the priority evaluation value EV O is the priority evaluation value EV of the person on the right (P[1]~P[6]) O In the vehicle V1 traveling left, more attention should be paid to avoiding contact with the person on the left than with the person on the right. Therefore, in the vehicle V1 traveling left, the position evaluation value EV POS is higher than the person on the right.

[0126] In the leftward moving state, depending on the size of the right and left characters, the priority evaluation value EV O is the priority evaluation value EV of the person on the right O However, if the person on the left and the person on the right are classified into the same size class in the left-moving state, the priority evaluation value EV O is the priority evaluation value EV of the person on the right O As a result, a higher priority is set for the person on the left than for the person on the right. In the example of FIG. 16, person P[7] corresponds to the person on the left, and person P[1] corresponds to the person on the right. For example, if people P[7] and P[1] are both classified as medium size in the leftward moving state, the priority evaluation value EV O [7] is the priority evaluation value EV of person P[1] O As a result, a higher priority is set for person P[7] than for person P[1].

[0127] In addition, when two or more people are located in a common grid area, a person who is classified into a relatively large size class among the two or more people is assigned a higher priority evaluation value EV than a person who is classified into a relatively small size class. O is derived and a high priority is set. This makes it possible to actually apply or more easily apply high-precision skeleton detection to people whose size merits the application of high-precision skeleton detection. For example, in FIG. 16, if person P[1] is classified as large size and person P[2] is classified as medium size, then "EV O [1]>EV O [2]”, so a higher priority is set for person P[1] than for person P[2].

[0128] The skeleton detection unit F4 detects the highest priority of the people P[1] to P[n] by selecting the highest priority of the people P[1] to P[n]. H The high-precision skeleton detection model F4_H is assigned to only a few people, and the low-precision skeleton detection model F4_L is assigned to the remaining people. H=5". For example, in FIG. 16, consider a case where the first to fifth priorities are set for persons P[1] to P[5] and the sixth to twelfth priorities are set for persons P[6] to P

[12] . In this case, the skeleton detection model F4_H is assigned to persons P[1] to P[5], and the skeleton detection model F4_L is assigned to persons P[6] to P

[12] .

[0129] Priority Evaluation Value (EV) O As described in the first embodiment, when it is not possible to distinguish between the fifth priority and the sixth priority based on the first index alone, the distinction may be made based on other indexes.

[0130] <<Third Example>> A third embodiment will now be described. Although the method of setting priorities according to the speed of the vehicle V1 and the method of setting priorities according to the path of the vehicle V1 have been described separately in the first and second embodiments, these setting methods can be implemented in combination. That is, the priority setting unit F3 may set the priority of each person by setting each grid priority value according to the speed of the vehicle V1 and the path of the vehicle V1 (whether the vehicle V1 is traveling straight, to the right, or to the left). When the vehicle V1 is traveling straight, each grid priority value may be set by the method shown in the first embodiment. When the vehicle V1 is traveling right or left, a difference in grid priority value may be set between the right grid area and the left grid area by the method shown in the second embodiment.

[0131] <<Fourth Example>> A fourth embodiment will be described. The time required to complete skeleton detection for one input image IN is called the required processing time. The time required from when the image data of one input image IN is acquired by the controller 11 until skeleton detection information for each person in the input image IN is output from the skeleton detection unit F4 corresponds to the required processing time. The in-vehicle device 10 is required to keep the required processing time for one input image IN within a specified time (this requirement is called the processing time requirement). In the first and second embodiments, in order to satisfy this processing time requirement, the high precision upper limit number MAX HHowever, if the required processing time can be kept within the specified time, the high precision upper limit number MAX H Setting is not required.

[0132] For example, let us consider a case CS4 in which a total of 27 people P[1] to P

[27] are detected from a single input image IN. Consider the following allocations using the first and second patterns. In the first pattern, five people are assigned to the skeleton detection model F4_H, and the remaining 22 people are assigned to the skeleton detection model F4_L. In the second pattern, seven people are assigned to the skeleton detection model F4_H, and no skeleton detection model is assigned to the remaining 20 people (i.e., skeleton detection is not performed for the remaining 20 people). In the second pattern, the controller 11 has more resources available because skeleton detection for the 20 people is not performed using the skeleton detection model F4_L. As a result, the overall processing load of the skeleton detection unit F4 in the first pattern can be made equal to or less than the overall processing load of the skeleton detection unit F4 in the second pattern, and the processing time requirement can be met in both the first and second patterns. Therefore, no problem occurs whether the skeleton detection unit F4 employs the first or second pattern.

[0133] Considering these circumstances, in the fourth embodiment, the high precision upper limit number MAX HInstead of determining the priority, the number of people assigned to the skeleton detection model F4_H is dynamically set for each input image IN. This setting may be performed by the priority setting unit F3 or the skeleton detection unit F4. As in the above-described embodiments, in the fourth embodiment, a high-precision skeleton detection model F4_H is preferentially assigned to people assigned a higher priority. In the above-described case CS4, the skeleton detection unit F4 can arbitrarily select either the first or second pattern. In the above-described case CS4, if the processing time requirement can be met even with the third pattern, the skeleton detection unit F4 may arbitrarily select either the first to third patterns. This selection may be made depending on the position and size of each person in the input image IN, the speed and steering angle of the vehicle V1, and the total number of people P[1] to P[n] detected from the input image IN. In the third pattern, six people assigned with the first to sixth priorities are assigned to the skeleton detection model F4_H, and ten people assigned with the seventh to sixteenth priorities are assigned to the skeleton detection model F4_L. In the third pattern, no skeleton detection model is assigned to the remaining 11 people (that is, skeleton detection is not performed on the remaining 11 people).

[0134] <<Fifth Example>> A fifth embodiment will be described. Fig. 17 shows an operation flowchart of the controller 11 in the in-vehicle device 10. Each process of steps S10 to S16 shown in Fig. 17 is executed by the controller 11. When the vehicle V1 starts, the controller 11 and the camera 51 start, and the operation of the controller 11 starts from the process of step S10. After the camera 51 starts, the camera 51 sequentially takes images at a predetermined shooting frame rate, and image data of an input image IN generated for each image is supplied to the in-vehicle device 10. In step S10, the controller 11 acquires image data of one input image IN.

[0135] In the following step S11, the object detection unit F1 generates object detection information by performing object detection processing based on the image data of the input image IN (see also FIG. 8). Thereafter, in step S12, the region division unit F2 generates region division information by performing region division processing based on the object detection information. However, when fixed region division processing is performed by the region division unit F2, the region division processing is performed without relying on the object detection information. In this case, the fixed region division processing may be performed before step S11 or simultaneously with step S11. After the object detection processing and region division processing, the process proceeds to step S13.

[0136] In step S13, the priority setting unit F3 executes a priority setting process based on the object detection information and the region segmentation information, thereby generating priority setting information. As a method for setting priorities in the priority setting process, any of the above-described methods (including the methods described in the first, second, or third embodiments) can be used, as well as the methods described in other embodiments described below. Then, in step S14, the skeleton detection unit F4 executes a skeleton detection process on the input image IN based on the image data of the input image IN while referring to the object detection information and the priority setting information, thereby generating skeleton detection information. In the subsequent step S15, the behavior prediction unit F5 executes a behavior prediction process. In the behavior prediction process, the behavior prediction unit F5 predicts the behavior of each person in the input image IN by detecting the posture of each person in the input image IN based on the skeleton detection information while referring to the object detection information, thereby generating behavior prediction information. Then, in step S16, the safety support unit F6 executes a safety support process to support the safe driving of the vehicle V1 based on the behavior prediction information. Then, the process returns to step S10, and the processes of steps S10 to S16 are repeated for a new input image IN.

[0137] Here, a note will be made regarding the technology of the in-vehicle device 10, with particular attention paid to priority setting and skeleton detection.

[0138] Skeleton detection enables posture detection and behavior prediction for each object (each person). The skeleton detection unit F4 performs skeleton detection for each object (each person) in the input image IN using multiple skeleton detection models, each with a different number of detected keypoints. However, the processing load of skeleton detection increases as the number of keypoints increases. For this reason, it is preferable to preferentially assign objects considered to be of high importance (objects considered to have a high need for posture detection and behavior prediction) to the skeleton detection model with the most keypoints (the skeleton detection model with the highest accuracy). On the other hand, the importance of each object (e.g., importance to ensuring safety when applied to a vehicle) can be estimated based on the image data of the input image IN. Therefore, the priority setting unit F3 sets a priority for each object based on the image data of the input image IN, and the skeleton detection unit F4 assigns each object (each person) to one of the multiple skeleton detection models based on the priority setting result. This allows objects considered to be of high importance to be preferentially assigned to the skeleton detection model with the most keypoints. As a result, it is possible to appropriately balance the required skeleton detection accuracy (required number of keypoints).

[0139] The multiple skeleton detection models include a first skeleton detection model with a relatively large number of key points and a second skeleton detection model with a relatively small number of key points. Skeleton detection models F4_H and F4_L (see FIG. 13) are examples of the first and second skeleton detection models. The skeleton detection unit F4 assigns objects with a relatively high priority among multiple objects in the input image IN to the first skeleton detection model preferentially over objects with a relatively low priority. This makes it possible to assign objects considered to be of high importance preferentially to the skeleton detection model with a large number of key points. As a result, it is possible to balance the processing load and the usefulness of the detection content (for example, usefulness for ensuring safety when applied to a vehicle).

[0140] The priority setting unit F3 sets a priority for each object based on the position of each object in the input image IN, the size of each object in the input image IN, and the driving state of the vehicle V1. Setting priorities by comprehensively taking this information into consideration enables appropriate priority setting according to various situations. In other words, setting priorities taking size into consideration makes it possible to apply an appropriate skeleton detection model according to size. Furthermore, the importance of skeleton detection for an object at a certain position in the input image IN may be relatively high under certain driving conditions, but relatively low under other driving conditions. Setting priorities taking the position of the object and the driving state of the vehicle V1 into consideration enables appropriate priority setting according to the driving scene.

[0141] In addition, the priority evaluation value EV O Priority can be set through the derivation of , and the priority evaluation value EV O The position evaluation value EV POS and size evaluation value EV SZ By including the above formula (1A) or formula (1B), the position and size of each object are taken into consideration (see above formula (1A) or formula (1B)). In this case, the grid priority value corresponding to the position of the object is the position evaluation value EV POS However, the grid priority value of each grid area is set depending on the running state of the vehicle V1 (see the first and second embodiments). This causes the priority of each object to vary depending on the running state of the vehicle V1.

[0142] Specifically, the priority setting unit F3 divides the input image IN into multiple grid areas and identifies in which grid area each object is located. Then, the priority setting unit F3 may set a priority for each object based on the result of this identification, the size of each object in the input image IN, and the speed of the vehicle V1 (see the first embodiment). This makes it possible to set an appropriate priority according to the speed of the vehicle V1. This is because the importance of skeleton detection for an object at a certain position in the input image IN varies depending on the speed of the vehicle V1.

[0143] More specifically, as explained in detail in the first embodiment, when comparing a near object (near person) and a distant object (distant person), a relatively high priority is likely to be set to the near object when the vehicle V1 is moving at a low speed, and to the distant object when the vehicle V1 is moving at a high speed. Which object actually has a higher priority depends on the size of each object in the input image IN. At least if the near object and the distant object are classified into the same size class, a relatively high priority is set to the near object when the vehicle V1 is moving at a low speed, and to the distant object when the vehicle V1 is moving at a high speed. This assigns a higher priority to objects that require more attention in relation to the speed of the vehicle V1.

[0144] In this case, when two or more objects are located in a common grid area, objects of a relatively large size class are given a higher priority than objects of a relatively small size class, so that the first skeleton detection model is actually applied or is more likely to be applied to objects having a size that merits application of the first skeleton detection model.

[0145] The priority setting unit F3 also divides the input image IN into multiple grid areas and identifies in which grid area each object is located. The priority setting unit F3 may then set a priority for each object based on the results of this identification, the size of each object in the input image IN, and information related to the steering angle of the vehicle V1 (see the second embodiment). Steering angle information or acceleration information (see FIG. 5) can be used as information related to the steering angle of the vehicle V1. This enables appropriate priority setting according to the path of the vehicle V1. This is because the importance of skeleton detection for an object at a certain position in the input image IN may be relatively high under certain path conditions, but relatively low under other path conditions.

[0146] More specifically, as explained in detail in the second embodiment, when comparing a right-side object (person on the right) and a left-side object (person on the left), a relatively high priority is likely to be set for the right-side object when the vehicle V1 is moving to the right, and for the left-side object when the vehicle V1 is moving to the left. Which object actually has a higher priority depends on the size of each object in the input image IN. At least if the right-side object and the left-side object are classified into the same size class, a relatively high priority is set for the right-side object when the vehicle V1 is moving to the right, and for the left-side object when the vehicle V1 is moving to the left. This assigns a higher priority to objects that require greater attention in relation to the path of the vehicle V1.

[0147] In this case, when two or more objects are located in a common grid area, objects of a relatively large size class are given a higher priority than objects of a relatively small size class, so that the first skeleton detection model is actually applied or is more likely to be applied to objects having a size that merits application of the first skeleton detection model.

[0148] Then, the behavior prediction unit F5 predicts the behavior of each object based on the detection results (skeleton detection information) of the key points of each object by the skeleton detection unit F4. The greater the number of key points, the higher the accuracy of behavior prediction, so it is possible to perform more accurate behavior prediction for objects that require greater attention. In other words, a better balance can be achieved between processing load and behavior prediction accuracy.

[0149] For example, consider a case where a large number of people appear in the input image IN just before vehicle V1 enters an intersection in a busy shopping district. In such a case, performing high-precision skeleton detection on all people would be unrealistic or difficult to meet the processing time requirements, as it would result in an excessively large processing load. Therefore, by taking into account the vehicle's driving conditions, etc., the high-precision skeleton detection model F4_H is assigned to people of higher importance. This makes it possible, for example, to assign the high-precision skeleton detection model F4_H only to people who are likely to suddenly appear in front of vehicle V1.

[0150] <<Sixth Example>> A sixth embodiment will be described. The controller 11 may be capable of executing a specific detection process for detecting the presence or absence of a maskable object based on image data of the input image IN. If the presence of a maskable object is detected, the specific detection process identifies the area in the input image IN where the image of the maskable object exists.

[0151] An occludable object is an object that may be located between the vehicle V1 and any person P[i] in real space, or more precisely, an object that may be located between the camera 51 and any person P[i]. Note that the body of the vehicle V1 (such as the hood) does not qualify as an occludable object. When an occludable object is located between the vehicle V1 and any person P[i] in real space, part or all of the person P[i] is occluded by the occludable object as viewed from the camera 51, and part or all of the image of the person P[i] is not included in the input image IN.

[0152] It is preferable that an object detection unit that performs specific detection processing is provided in the controller 11 separately from the object detection unit F1. In the specific detection processing, a specific type of object is detected as an object that can be obscured. The specific type of object is a vehicle or a utility pole, etc., and is different from the object to be detected by the object detection unit F1. The vehicle detected in the specific detection processing is a vehicle different from vehicle V1, for example, a vehicle located diagonally in front of vehicle V1.

[0153] FIG. 18 shows an example of an input image IN containing an image of a occludable object. Images of n people appear in the input image IN, but in FIG. 18, only the symbols (P[1], P[2]) for people P[1] and P[2] are explicitly indicated. In FIG. 18, a bus 710 is an occludable object. In real space, the bus 710 is located diagonally forward and to the left of vehicle V1. In the input image IN of FIG. 18, the image of the bus 710 exists across a total of four grid regions R[3,5], R[4,5], R[3,6], and R[4,6]. In the input image IN, the grid region in which the image of the occludable object is located is referred to as a specific grid region. In the example of FIG. 18, only grid regions R[3,5], R[4,5], R[3,6], and R[4,6] correspond to specific grid regions. The priority setting unit F3 sets a priority evaluation value EV of a person of interest detected in the object detection process when the person of interest is located in a specific grid area, compared to when the person of interest is not located in a specific grid area. O Increase.

[0154] For example, the priority setting unit F3 sets the priority evaluation value EV O [i] can be derived according to the following formula (2A) or formula (2B): POS [i] and size evaluation value EV SZ The method for deriving [i] is as described above. J[i] is an adjustment value. In principle, the adjustment value J[i] is zero. A positive predetermined value is set to the adjustment value J[i] only when the grid area in which person P[i] is located corresponds to a specific grid area. EV O [i]=EV POS [i]×EV SZ [i]+J[i] (2A) EV O [i]=EV POS [i]+EV SZ [i]+J[i] (2B)

[0155] In the example of FIG. 18, person P[1] is located in grid area R[4,6], which corresponds to the specific grid area, so a positive predetermined value is set to the adjustment value J[1]. In the example of FIG. 18, person P[2] is located in grid area R[7,7], which does not correspond to the specific grid area, so zero is set to the adjustment value J[2]. As a result, the priority evaluation value EV O If the predetermined value set for the adjustment value J[1] is set high enough, the highly accurate skeleton detection model F4_H can be reliably assigned to person P[1].

[0156] A person located in a specific grid area (person P[1] in the example of FIG. 18) may be difficult for user U1 to recognize due to the influence of obscuring objects. In the example of FIG. 18, person P[1] may have appeared from a location in the shadow of bus 710, and particular attention should be paid to such a person. According to the method of this embodiment, a highly accurate skeleton detection model F4_H is more likely to be assigned to a person located in a location susceptible to the influence of obscuring objects (person P[1] in the example of FIG. 18). As a result, it becomes easier to perform detailed behavior predictions of people who are difficult for user U1 to recognize, thereby promoting safety assurance through safety support processing.

[0157] <<Seventh Example>> A seventh embodiment will be described. Based on the method shown in the second embodiment, in a rightward or leftward moving state, a lower priority may be set to a person located in the center grid area compared to people located in the right grid area and the left grid area. To achieve this, in a rightward or leftward moving state, the priority setting unit F3 may set the grid priority value of the center grid area lower than the grid priority values ​​of the right grid area and the left grid area.

[0158] In the rightward traveling state, the driver often pays attention to pedestrians and the like on the right side, but is also likely to pay a fair amount of attention to people located directly in front of the vehicle V1. Therefore, in the rightward traveling state, the driver may be less aware of people who suddenly appear in front of the vehicle V1 from the left side. Therefore, in the rightward traveling state, it is advisable to set the grid priority value of the right grid area higher than the grid priority value of the left grid area, and to set the grid priority value of the left grid area higher than the grid priority value of the center grid area. As a result, in the rightward traveling state, a certain degree of priority is set for people on the left, making it possible to take care of people who suddenly appear in front of the vehicle V1 from the left side. Conversely, in the leftward traveling state, it is advisable to set the grid priority value of the left grid area higher than the grid priority value of the right grid area, and to set the grid priority value of the right grid area higher than the grid priority value of the center grid area.

[0159] <<Eighth Example>> An eighth embodiment will be described. The region division unit F2 may change the content of the region division process according to the object detection information. For example, the region division unit F2 may change the values ​​of M and N according to the total number of people P[1] to P[n] detected from the input image IN by the object detection process (i.e., the value of n).

[0160] The region dividing unit F2 may also mix multiple sizes for the grid regions R[1,1] to R[M,N] depending on the distribution of people in the input image IN. That is, for example, depending on the distribution of people in the input image IN, grid regions R[1,1] to R[M,N] may include a mixture of grid regions having a first size and grid regions having a second size. Here, the first size is smaller than the second size. Grid regions having a first size are referred to as small grid regions, and grid regions having a second size are referred to as large grid regions. For example, if people are densely concentrated in a portion of the input image IN, small grid regions are set in the portion where people are densely concentrated, and large grid regions are set in the other portions. More specifically, for example, the region dividing unit F2 first sets multiple large grid regions in the input image IN by dividing the entire image region of the input image IN at equal intervals in each of the X-axis and Y-axis directions. Thereafter, the area division unit F2 counts the total number of people located in each large grid area based on the object detection information, and identifies a large grid area where the total count is equal to or greater than a predetermined number. The area division unit F2 then divides the identified large grid area into multiple areas in each of the X-axis and Y-axis directions, thereby setting multiple small grid areas in the identified large grid area.

[0161] <<Ninth Example>> A ninth embodiment will now be described.

[0162] The priority setting unit F3 may always assign the lowest priority to persons classified as extremely small, which makes it difficult to assign a high-precision skeleton detection model F4_H to persons classified as extremely small. This is because even if a high-precision skeleton detection model F4_H is applied to a person classified as extremely small, it is difficult to accurately detect each keypoint due to their small size. Furthermore, the posture, etc. of a person who is sufficiently small is likely to be of little importance for behavior prediction or safety support. When person P[i] is classified as extremely small, the priority setting unit F3 may assign the lowest priority to person P[i] and include a very small size flag of “1” associated with person P[i] in the priority setting information. When person P[i] is classified as large, medium, or small, a very small size flag of “1” associated with person P[i] is not included in the priority setting information. When a minimum size flag of "1" is associated with person P[i], the skeleton detection unit F4 may always assign a low-accuracy skeleton detection model F4_L to person P[i] or may not perform skeleton detection on person P[i].

[0163] The priority setting unit F3 may switch the priority setting method depending on the type of road on which the vehicle V1 is currently traveling. That is, for example, the controller 11 (e.g., the priority setting unit F3) determines whether the road on which the vehicle V1 is currently traveling is an expressway or an ordinary road other than an expressway. The controller 11 may make this determination based on map information indicating the road type at each point and vehicle position information. Alternatively, the controller 11 may make this determination by image recognition based on image data of the input image IN (e.g., image recognition of road signs indicating the road's speed limit). If the road on which the vehicle V1 is currently traveling is an expressway, the priority setting unit F3 may execute the same priority setting method as for the high-speed state in the first embodiment. If the road on which the vehicle V1 is currently traveling is an ordinary road, the priority setting unit F3 may execute the same priority setting method as for the low-speed state in the first embodiment.

[0164] Consider a case where a low-accuracy skeleton detection model F4_L is assigned to a certain person of interest in a first input image IN, and the behavior prediction process predicts that the person of interest is likely to appear on the predicted driving path of vehicle V1. In this case, the priority setting unit F3 may assign a sufficiently high priority to the person of interest in a second input image IN, so that a high-accuracy skeleton detection model F4_H is assigned to the person of interest. The second input image IN is an image captured by camera 51 after the first input image IN. The controller 11 tracks and identifies the person of interest in a video sequence made up of multiple input images IN, including the first input image IN and the second input image IN, based on the image data of the multiple input images IN.

[0165] When a certain skeleton detection model (F4_H or F4_L) is assigned to a person of interest, the skeleton detection unit F4 may dynamically change the specific body parts detected by the skeleton detection model according to various situations, such as allocating more key points to the feet or the upper body depending on the situation.

[0166] While the method for assigning each person to a skeleton detection model has been described with a focus on a high-precision skeleton detection model (F4_H) and a low-precision skeleton detection model (F4_L), the skeleton detection unit F4 may have three or more skeleton detection models. Even when the skeleton detection unit F4 is provided with three or more skeleton detection models, each person may be assigned to a skeleton detection model according to priority. For example, if "n=20" and first to twentieth priorities are assigned to persons P[1] to P

[20] , a high-precision skeleton detection model may be assigned to persons P[1] to P[5], and a standard-precision skeleton detection model may be assigned to persons P[6] to P

[10] . In this case, a low-precision skeleton detection model may be assigned to persons P

[11] to P

[20] . Alternatively, a low-precision skeleton detection model may be assigned to persons P

[11] to P

[15] , and skeleton detection for persons P

[16] to P

[20] may not be performed.

[0167] Although the above embodiments have been described mainly assuming that the object to be detected by the object detection unit F1 is a person, the present invention can also be applied to cases where the object to be detected is an animal other than a person.

[0168] A program that causes a computer device to execute any of the methods described in the embodiments of the present invention, and a non-volatile recording medium on which the program is recorded, are included within the scope of the embodiments of the present invention. The program that causes a computer device to execute any of the methods described in the embodiments of the present invention may be a subprogram incorporated into any main program or called by any main program. Any processing in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software equivalent to the program, or a combination of hardware and software.

[0169] The in-vehicle device 10 or the controller 11 is a type of computer device. The functional blocks F1 to F4 shown in FIG. 8 constitute a skeleton detection device, and the functional blocks F1 to F5 shown in FIG. 8 constitute a behavior prediction device. The methods performed by the skeleton detection device and the behavior prediction device can be referred to as a skeleton detection method and a behavior prediction method, respectively. The programs that cause the computer device to perform the skeleton detection method and the behavior prediction method can be referred to as a skeleton detection program and a behavior prediction program, respectively. By executing the skeleton detection program and the behavior prediction program on the computer device, multiple skeleton detection models are formed in the computer device. [Explanation of symbols]

[0170] 1. In-vehicle systems V1 vehicle U1 User ST1 seat 10 Onboard equipment 11 Controller 12 Memory 13 Communications Department 14 Recording media 20 Driving control device 30 Actuator section 40 Vehicle sensor unit 50 Camera Department 51 Camera 60 HMI 61 Display device 62 Speaker 63 Operation input section IN Input image R[x,y] Grid area F1 Object detection unit F2 area division part F3 Priority setting section F4 Skeleton detection unit F4_H Skeleton detection model (high accuracy) F4_L Skeleton detection model (low accuracy) F5 Behavior Prediction Department F6 Safety Support Department

Claims

1. A skeleton detection device has a plurality of skeleton detection models each having a different number of detected key points, and detects key points of each object in an input image using the plurality of skeleton detection models based on image data of the input image including images of the objects, A priority is set for each object based on the image data of the input image, and each object is assigned to one of the plurality of skeleton detection models based on the priority setting result. ,Skeletal detection device.

2. the plurality of skeleton detection models include a first skeleton detection model and a second skeleton detection model, the number of key points detected by the first skeleton detection model being greater than the number of key points detected by the second skeleton detection model; Among the plurality of objects, an object to which a relatively high priority is set is assigned to the first skeleton detection model in preference to an object to which a relatively low priority is set. The skeletal detection device according to claim 1 .

3. An image captured by a camera installed in a vehicle is acquired as the input image; The priority is set for each object based on the position of each object in the input image, the size of each object in the input image, and the running state of the vehicle. The skeletal structure detection device according to claim 1 or 2.

4. Dividing the input image into a plurality of grid areas and identifying in which grid area each object is located; The priority is set for each object based on the result of the identification, the size of each object in the input image, and the speed of the vehicle. The skeletal structure detection device according to claim 3 .

5. classifying the size of each object in the input image into a plurality of size classes; When the size of a near object located in a near grid area and the size of a far object located in a far grid area are classified into a common size class, if the speed of the vehicle falls within a first speed range, a relatively higher priority is set for the near object than for the far object, and if the speed of the vehicle falls within a second speed range that is higher than the first speed range, a relatively higher priority is set for the far object than for the near object; the near grid region and the far grid region are any of the plurality of grid regions, The near grid area is a grid area in which an image of an object that is relatively closer to the vehicle than the far grid area appears. The skeletal structure detecting device according to claim 4 .

6. When two or more objects are located in a common grid area, the priority is set higher for an object classified into a relatively large size class among the two or more objects than for an object classified into a relatively small size class. The skeletal detection device according to claim 5 .

7. Dividing the input image into a plurality of grid areas and identifying in which grid area each object is located; The priority is set for each object based on the result of the identification, the size of each object in the input image, and steering angle information of the vehicle or acceleration information of the vehicle that reflects the steering angle of the vehicle. The skeletal structure detecting device according to claim 3 .

8. determining whether the vehicle is in a rightward traveling state turning right or a leftward traveling state turning left based on the steering angle information or the acceleration information of the vehicle; classifying the size of each object in the input image into a plurality of size classes; When the size of a right object located in the right grid area and the size of a left object located in the left grid area are classified into a common size class, if it is determined that the vehicle is in the rightward traveling state, a relatively higher priority is set for the right object than for the left object, and if it is determined that the vehicle is in the leftward traveling state, a relatively higher priority is set for the left object than for the right object; the right grid area and the left grid area are any of the plurality of grid areas, The right grid area is a grid area in which an image of a right area outside the vehicle appears, and the left grid area is a grid area in which an image of a left area outside the vehicle appears. The skeletal structure detection device according to claim 7 .

9. When two or more objects are located in a common grid area, the priority is set higher for an object classified into a relatively large size class among the two or more objects than for an object classified into a relatively small size class. The skeletal detection device according to claim 8 .

10. 3. A method for predicting the behavior of a person or an animal as each object based on the results of detecting key points of each object, the method comprising the skeleton detection device according to claim 1 or 2. ,Behavioral prediction device.

11. A vehicle equipped with a behavior prediction device that includes the skeleton detection device according to claim 1 or 2 and predicts the behavior of each object, such as a person or an animal, based on the detection results of key points of each object, The image captured by a camera installed in the vehicle is acquired by the skeleton detection device as the input image. ,vehicle.

12. A computer device is configured to generate a plurality of skeleton detection models each having a different number of detected key points; Detecting key points of each object in an input image using the plurality of skeleton detection models based on image data of the input image including images of the plurality of objects; setting a priority for each object based on image data of the input image, and assigning each object to one of the plurality of skeleton detection models based on the priority setting result. , a skeleton detection program.

Citation Information

Patent Citations

  • Skelton detection system

    JP2022042233A