Method and device for detecting children
By acquiring images from cameras and generating bird's-eye view images, extracting pixel coordinates of human feet and head, and estimating distance and height, the problem of distinguishing between adult and child control in autonomous driving systems has been solved, improving safety and user-friendliness.
Patent Information
- Application Number
- CN202411752535.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-10
- Filing Date
- 2024-12-02
- Publication Date
- 2025-11-11
Smart Images

Figure CN120931718A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to Korean Patent Application No. 10-2024-0061714, filed with the Korean Intellectual Property Office on May 10, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to a method and apparatus for detecting children. Background Technology
[0004] In order to move in spaces where people exist, mobile systems (e.g., vehicles, robots) can be designed with human interaction, safety, user-friendliness, and other considerations in mind. For example, since mobile systems such as self-driving robots, self-driving vehicles, warehouse robots, and / or self-driving pallets, which include delivery robots and / or cleaning robots, move in spaces where people live / work / exist, it is important to prevent collisions with people.
[0005] The matters described in this background section are only intended to enhance the understanding of the background of this disclosure and should not be construed as an admission that these matters correspond to prior art known to those skilled in the art. Summary of the Invention
[0006] The following description of the invention presents a brief overview of some features. This description is not a broad overview and is not intended to identify key or important elements.
[0007] Systems, apparatus, and methods for controlling a vehicle (e.g., based on the detection of a child and / or the height of a detected person) are described. A method for controlling the operation of a vehicle may include: acquiring an image from a camera on the vehicle; extracting first foot pixel coordinates corresponding to the feet of a detected person in the acquired image and head pixel coordinates corresponding to the person's head in the acquired image, based on semantic segmentation of the image; generating a transformed view image by converting the coordinate system of the acquired image to a transformed view coordinate system; determining second foot pixel coordinates corresponding to the first foot pixel coordinates in the transformed view coordinate system based on the transformed view image; estimating the distance between the vehicle and the detected person based on the first foot pixel coordinates; estimating the height of the detected person based on the distance and the head pixel coordinates; and controlling the operation of the vehicle based on the estimated height.
[0008] An apparatus for a vehicle may include one or more processors configured to execute program code loaded on one or more memory devices. When executed by the one or more processors, the program code may configure the apparatus to: acquire images from a camera of the vehicle; extract, based on semantic segmentation of the images, first foot pixel coordinates corresponding to the feet of a detected person in the image and head pixel coordinates corresponding to the person's head; generate a transformed view image by converting the coordinate system of the acquired image to a transformed view coordinate system; determine, based on the transformed view image, second foot pixel coordinates corresponding to the first foot pixel coordinates in the transformed view coordinate system; estimate the distance between the vehicle and the detected person based on the first foot pixel coordinates; estimate the height of the detected person based on the distance and the head pixel coordinates; and control the operation of the vehicle based on the estimated height.
[0009] These and other features and advantages are described in more detail below. Attached Figure Description
[0010] The above and other objects, features and other advantages of this disclosure will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings, wherein:
[0011] Figure 1 This is a block diagram illustrating a device for detecting children, based on an example.
[0012] Figure 2 This is a flowchart illustrating a method for detecting children based on an example.
[0013] Figure 3 Examples of devices and methods for detecting children are shown.
[0014] Figure 4 Examples of devices and methods for detecting children are shown.
[0015] Figure 5 and Figure 6 Examples of devices and methods for detecting children are shown.
[0016] Figure 7 Examples of devices and methods for detecting children are shown.
[0017] Figure 8 Examples of devices and methods for detecting children are shown.
[0018] Figure 9 Examples of devices and methods for detecting children are shown.
[0019] Figure 10 This is a diagram illustrating a computing device according to an example. Detailed Implementation
[0020] Referring to the accompanying drawings, examples of this disclosure will be described in detail below to enable those skilled in the art to readily implement it. However, this disclosure may be implemented in many different forms and is not limited to the examples described herein. For clarity in the accompanying drawings, portions irrelevant to the description have been omitted, and throughout the specification, the same reference numerals denote the same elements. Unless otherwise defined, terms used herein, including technical or scientific terms, may have the meaning commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] Throughout the specification and claims, unless expressly stated otherwise, the word "comprising" and variations such as "including" or "comprises" are to be understood as implying inclusion of the stated element but not excluding any other element. Terms including designations such as first, second, etc., may be used to describe a variety of elements, but these elements are not limited by the terms. These terms are used only for the purpose of distinguishing one element from another. Unless the context otherwise indicates, singular expressions used herein may have a plural meaning, and this also applies to singular expressions described in the claims.
[0022] Terms such as “component,” “part,” “group,” “module,” “unit,” and “method” used in this specification may refer to a unit / entity that processes / performs at least one function or operation described in this specification, and the unit / entity may be implemented as hardware or software, or a combination of hardware and software. Additionally, at least some components or functions of the devices and methods for detecting children according to the examples described below may be implemented as programs or software, which may be stored on a computer-readable medium. In particular, these terms generally refer to items of functions or groups that can be logically grouped together to perform related functions. The same reference numerals are generally intended to refer to the same or similar components. Components, units, and modules may be implemented as software, hardware, or a combination of software and hardware. The aforementioned components, units, modules, and / or functions may be implemented and / or performed by one or more processors. For example, components, units, and / or modules may include processors, microprocessors, graphics processing units, logic circuits, application-specific circuits, application-specific integrated circuits, programmable array logic, field-programmable gate arrays, controllers, microcontrollers, and / or other suitable hardware. Components, units, and / or modules may also include software control modules implemented using, for example, processors or logic circuits. Components, units, and / or modules may include or additionally have access to memories such as: for example, one or more non-transitory computer-readable storage media such as: random access memory, read-only memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, flash memory / other memory devices, data registers, databases, and / or other suitable hardware. One or more types of storage media may include any or all of the following tangible memories or related modules of a computer, processor, etc., that can provide non-transitory storage at any time for software programming: various semiconductor memories, magnetic tape drives, disk drives, etc.
[0023] As used herein, when describing a component (e.g., a first component) as “connected” or “linked” to another component (e.g., a second component), this may mean not only that the component is directly connected to or linked to the other component, but also that the component is connected to or linked to another component through a third component (e.g., a third component).
[0024] The expression “based on” as used in this document is intended to describe one or more factors that influence the behavior or operation of a determination or decision as described in the phrase or sentence that includes the expression, and the expression does not exclude any other factors that influence the behavior or operation of a determination or decision.
[0025] The exemplary phrases “at least one of A; B; or C” or “at least one of A, B, or C” used for the purposes of this application and claims mean at least one A, or at least one B, or at least one C, or any combination of at least one A, at least one B, and at least one C. Furthermore, exemplary phrases such as “A, B, and C,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” as used herein may refer to each of the listed items or all possible combinations of the listed items. For example, “at least one of A or B” may mean (1) at least one A; (2) at least one B; or (3) at least one A and at least one B.
[0026] Given the differences in behavioral patterns between adults and children, it may be useful to detect / differentiate adults and children separately when preventing / avoiding collisions and to control the vehicle / mobile system differently to avoid adults / children and / or keep them safe, based on the detection / differentiation. For example, relative to adults, the vehicle / mobile system could be controlled to move slower when bypassing / avoiding children, avoid children by a greater distance, stop and / or wait longer to allow children to pass / move, and stop further away from pedestrian areas (e.g., crosswalks) if a child is detected nearby (e.g., within a threshold distance). On the other hand, for the control and / or remote control of the mobile system, generating a bird's-eye view image may be beneficial, allowing for an immediate 360-degree inspection of the surrounding situation. By distinguishing adults and children from each other in the bird's-eye view image, collision avoidance strategies or other operations for autonomous driving control can be subdivided based on whether the detected person is an adult or a child.
[0027] Operational control for automated driving of a mobile system / vehicle can include various driving controls of the mobile system / vehicle by means of the devices disclosed herein (e.g., vehicle control devices). For example, various driving controls can include acceleration, deceleration, steering control, gear shifting control, braking system control, traction control, stability control, cruise control, lane keeping assist control, collision avoidance system control, emergency braking assist control, traffic sign recognition control, adaptive headlight control, etc. Different controls / control settings can be applied depending on whether the detected person is an adult or a child.
[0028] A bird's-eye view can be a view taken from a distance above the ground and / or an object, and can capture an area larger than a threshold (e.g., a threshold region configured in the aircraft's memory). The bird's-eye view image can indicate (and / or can be associated with) the perspective angle from the aircraft (e.g., line, yaw, and pitch information from the aircraft and / or one or more cameras of the aircraft). For example, a bird's-eye view image can be generated by, for example, transforming the coordinate system of the bird's-eye view image by converting a first coordinate system of the camera's field of view to a transformed coordinate system. The bird's-eye view image can indicate (and / or can be associated with) the temporal information and / or other indicators of the frames of the bird's-eye view image. The bird's-eye view image can indicate (and / or can be associated with) one or more landmark images contained within the bird's-eye view image.
[0029] Figure 1 This is a block diagram illustrating a device for detecting children, based on an example.
[0030] Reference Figure 1 According to the example, the child detection device 10 can execute one or more instructions (e.g., program code) loaded / stored on one or more memory devices via one or more processors. For example, the child detection device 10 can be implemented as described below. Figure 10 The computing device 50 is described above. In this case, one or more processors may correspond to processor 510 of computing device 50, and one or more memory devices may correspond to memory 530 of computing device 50. One or more instructions (e.g., program code) may be executed by one or more processors to perform the function of detecting children. One or more processors may be part of a mobile system (e.g., a vehicle) configured to travel in spaces where people may be present.
[0031] The device 10 for detecting children, according to the example, may include an image acquisition module 110, a person detection module 120, a bird's-eye view image generation module 130, a person distance estimation module 140, a person height estimation module 150, a child identification module 160, and a display module 170.
[0032] Image acquisition module 110 can acquire images from a camera installed in a mobile system that enables mobility. Here, a mobile system can be a system that moves in spaces where people may live / work / exist. For example, a mobile system can be a self-driving robot, self-driving vehicle, warehouse robot, and / or self-driving pallet, such as a delivery robot or cleaning robot. The scope of mobile systems in this document is not limited to the examples listed.
[0033] In the examples, the mobile system may include multiple monocular cameras. One or more of the monocular cameras (e.g., each) can acquire / generate / obtain images by utilizing a single optical lens system (e.g., through photography). Monocular cameras are advantageous (e.g., unlike stereo cameras) because they are uncomplicated, inexpensive, lightweight, and easy to use in various environments. In some examples, the mobile system may include multiple monocular cameras (e.g., two, three, four monocular cameras, etc.). Images acquired (e.g., captured) by one or more monocular cameras can be used to generate a bird's-eye view image of the area surrounding the mobile system (e.g., the area captured within the field of view of one or more cameras).
[0034] Each of one or more monocular cameras may have a corresponding field of view (FoV). FoV can refer to the observable area that a camera (or other vision sensor) can capture at any given moment. FoV can typically be measured as an angle, representing the range of a scene that can be seen horizontally, vertically, or diagonally. In the case of cameras and sensors, a wider FoV can allow more of the surrounding environment to be captured in a single image or scan, which can be useful for applications such as photography, video recording, virtual reality, and / or autonomous navigation. The FoV of a given camera can be affected by a variety of factors such as lens design, sensor size, and / or the distance between the observer or device and the observed object. (In the case of one or more monocular cameras comprising multiple cameras) different cameras among the multiple cameras may have different (e.g., distinct and / or partially overlapping) fields of view (FoVs). Images acquired by multiple monocular cameras (e.g., still images and / or video images) can be combined (e.g., stitched together) to form a single image with a combined FoV that includes the FoV of each camera in the combination. The images in the image acquisition module 110 discussed herein may be images acquired by a single camera among one or more cameras, or (for example, in the case where one or more monocular cameras are multiple monocular cameras) combined images formed by corresponding images acquired by multiple monocular cameras.
[0035] The person detection module 120 can detect one or more people in an image acquired by the image acquisition module 110. In some examples, the person detection module 120 may include / perform object recognition and / or classification algorithms. In some examples, the person detection module 120 may employ machine learning models such as convolutional neural networks (CNNs), region-based CNNs (R-CNNs), or transfer learning training models to detect people in an image.
[0036] The human detection module 120 can detect pixels corresponding to a person's feet in each (e.g., at least one) of one or more images acquired by the image acquisition module 110 and in which a person is identified, by utilizing semantic segmentation, and extract the first foot pixel coordinate p of the corresponding pixel. foot The first pixel coordinate is shown below. foot It can include the x-coordinate indicating the position of the image on the x-axis. foot and the y-coordinate indicating the position of the image on the y-axis. foot .
[0037] p foot =(x foot ,y foot )
[0038] The human detection module 120 can extract pixels corresponding to a person's feet in each (e.g., at least one) of one or more images, detect pixels corresponding to a person's head by utilizing semantic segmentation, and extract the head pixel coordinates p of the corresponding pixel. head The following is the header pixel coordinate p. head This can include the x-coordinate. head and y coordinates head .
[0039] p head =(x head ,y head )
[0040] Semantic segmentation makes it possible to classify each pixel in an image into which category it belongs. Labels can be assigned to each pixel constituting an image via semantic segmentation. Therefore, various objects (and / or constituent pixels) in an image can be accurately classified. The person detection module 120 can identify, at the pixel level, shapes and / or boundaries corresponding to / indicating people, feet, and / or heads in an image via semantic segmentation. For example, in the process of extracting features from an image using a deep learning model such as a CNN, a variety of information can be extracted, ranging from low-level features such as image edges, colors, and / or textures to high-level features such as the shape of objects and relationships between objects. Based on the extracted features, the neural network can classify each pixel of the image into one of several predefined categories to generate a segmentation map. The segmentation map can be output (e.g., where color codes are specified according to one or more categories).
[0041] In some examples, the first foot pixel coordinate p footThe coordinates of a specific pixel in a segmented region representing a person's foot can be set / selected. For example, the segmented region representing a person's foot can be set / selected as a first temporary region. The first temporary region can be a region corresponding to a person's foot detected in an image acquired by the image acquisition module 110. The coordinates of the bottommost pixel among one or more pixels contained in the first temporary region can be extracted as the first foot pixel coordinates p. foot Alternatively, another pixel, such as the center pixel (e.g., the center of mass), can be selected as the first foot pixel coordinate p. foot .
[0042] In some examples, the head pixel coordinates p head The coordinates can be set to the coordinates of a specific pixel in a segmented region representing a human head. For example, the segmented region representing a human head can be set as a second temporary region. The second temporary region can be a region corresponding to a human head detected in an image acquired by the image acquisition module 110. The coordinates of the topmost pixel among one or more pixels contained in the second temporary region can be extracted as the head pixel coordinates p. head Alternatively, another pixel, such as the center (e.g., the center of mass) pixel, can be selected as the first head pixel coordinate p. head .
[0043] The bird's-eye view image generation module 130 can generate a bird's-eye view image from or based on an image acquired by the image acquisition module 110. The bird's-eye view image can represent the objects or environment around the mobile system from a bird's-eye view perspective (e.g., from the sky, from the ceiling, etc.), thereby allowing for a clear identification of the overall structure and / or arrangement of objects above the ground area.
[0044] In some examples, the bird's-eye view image generation module 130 can acquire intrinsic and / or extrinsic parameters of the camera to generate a bird's-eye view image. Intrinsic parameters can be camera-specific information and can represent characteristics of the camera's lens and / or image sensor. For example, intrinsic parameters may include the focal length corresponding to the distance from the center of the lens to the image sensor, the principal point where the optical axis intersects the image plane on the image sensor, distortion coefficients for calibrating geometric and optical distortions of the lens, scaling factors representing the influence of the image sensor's pixel size on actual distance units, etc. Extrinsic parameters can represent the camera's position and / or orientation, and / or may include, for example, a rotation matrix representing the camera's orientation in three-dimensional (3D) space, a vector representing the camera's position, etc.
[0045] The bird's-eye view image generation module 130 can transform each point on the image acquired by the image acquisition module 110 into a point on the bird's-eye view image by utilizing / based on a homography matrix. A homography matrix is a transformation matrix that can be used to transform (and / or project) an image in one (first) plane to another (second) plane. For example, a homography matrix can be a 3×3 transformation matrix. If distortion occurs during the process of transforming each point on the image acquired by the image acquisition module 110 into a transformed point on the bird's-eye view image by utilizing the homography matrix, calibration can be performed based on (e.g., by utilizing) the intrinsic parameters of the camera.
[0046] The bird's-eye view image generation module 130 can generate a lookup table based on the result of the transformation. The bird's-eye view image generation module 130 can generate a bird's-eye view image based on (e.g., by utilizing) the lookup table, using an image acquired by the image acquisition module 110 (e.g., in real-time). The lookup table can store a previously calculated mapping from the position of each pixel in the original image (e.g., the image acquired by the image acquisition module 110) to the corresponding position on the bird's-eye view image. Therefore, a bird's-eye view can be generated quickly during real-time image processing.
[0047] The human distance estimation module 140 can be based on the first foot pixel coordinate p foot Estimate the distance d between the mobile system and the person detected in the image acquired by the image acquisition module 110. person Specifically, the human distance estimation module 140 can set the first foot pixel coordinate p in the normalized coordinate system. foot And the camera coordinates, set the coordinates in the world coordinate system (here, optionally, the bird's-eye view coordinate system) to correspond to the coordinates set in the normalized coordinate system, and then, based on the response to the first foot pixel coordinate p foot The coordinates of the camera are set in the world coordinate system to estimate the distance d between the mobile system and the person detected in the image acquired by the image acquisition module 110. person .
[0048] Human height estimation module 150 can be based on the distance d estimated by human distance estimation module 140. person and the head pixel coordinates p extracted by the human detection module 120 head Estimate the height h of the person detected in the image acquired by the image acquisition module 110. person .
[0049] If the height h is estimated by the height estimation module 150 personIf the value is within a predetermined range (e.g., meeting a child height standard, equal to or below a threshold, below a threshold), then the child determination module 160 can determine that the person detected in the image acquired by the image acquisition module 110 is a child. Optionally, if the height h estimated by the human height estimation module 150 is within a predetermined range (e.g., meeting a child height standard, equal to or below a threshold, below a threshold), then the child determination module 160 can determine that the person detected in the image acquired by the image acquisition module 110 is a child. person If the value is not within the range (e.g., does not meet the child height standard, is greater than the threshold, or is greater than or equal to the threshold), then the child determination module 160 can determine that the person detected in the image acquired by the image acquisition module 110 is not a child.
[0050] The bird's-eye view image generation module 130 can acquire / generate images with first foot pixel coordinates p from a bird's-eye view image (e.g., in a bird's-eye view and / or world coordinate system). foot The corresponding second pixel coordinate p' foot In some examples, the bird's-eye view image generation module 130 can obtain / generate the image with the first foot pixel coordinates p using a homography matrix. foot The corresponding second pixel coordinate p' foot .
[0051] Display module 170 can display the second-foot pixel coordinates p' obtained by bird's-eye view image generation module 130 on the bird's-eye view image. foot The child identification module 160 determines whether a person detected in an image acquired by the image acquisition module 110 is a child.
[0052] According to this example, the distance to a person can be accurately measured using only a monocular camera, without the use of expensive equipment such as LiDAR or RGB-D cameras. The distance to a person can be used to estimate their height, which can then be used to identify whether the person is an adult or a child. Classifying a person as an adult or a child can be used to control the mobility system differently to prevent collisions with people (e.g., when the mobility system is autonomous and / or driving through traffic control). Specifically, by detecting people in a bird's-eye view image and distinguishing between adults and children, the collision avoidance strategy of the mobility system can be refined / selected accordingly, and customized services and content related to mobility can be provided.
[0053] For convenience, examples will be used to describe this. Figures 2 to 5 and Figures 7 to 10 The steps in these examples are executed by the processor circuitry. Figures 2 to 5 and Figures 7 to 10 One, some, or all of the steps or a portion thereof of the example method may be performed by one or more other circuits. Figures 2 to 5 and Figures 7 to 10 One or more steps of the example method may be omitted, performed in a different order, and / or modified in other ways, and / or one or more additional steps may be added.
[0054] Figure 2 This is a flowchart illustrating a method for detecting children based on an example.
[0055] Reference Figure 2 The method for detecting children according to the example may include: extracting foot pixel coordinates corresponding to a person's feet and head pixel coordinates corresponding to a person's head from images acquired from one or more cameras (e.g., one or more monocular cameras) set in a mobile system S201; generating a bird's-eye view image from the acquired images S202; obtaining second foot pixel coordinates corresponding to the first foot pixel coordinates from the bird's-eye view image S203; estimating the distance between the mobile system and the detected person based on the first foot pixel coordinates S204; and estimating the height of the detected person based on the distance and head pixel coordinates S205.
[0056] For more detailed information on methods for detecting children, please refer to the description of the examples in the instruction manual; therefore, redundant descriptions are omitted here (e.g., see the description of the steps described herein).
[0057] Figure 3 Examples of devices and methods for detecting children are shown.
[0058] Reference Figure 3 Embodiments of devices and methods for detecting children may include: receiving RGB images from one or more cameras included in a mobile system in step S301; performing semantic segmentation in step S302; and detecting human and / or foot pixels in step S303. Embodiments may include: selecting pixel coordinates of the foot and head based on the segmentation map obtained through semantic segmentation (in steps S302 and / or S303) in step S304; and extracting the first foot pixel coordinate p in step S305. foot ; and in step S306, extract the head pixel coordinates p head .
[0059] Figure 4 Examples of devices and methods for detecting children are shown.
[0060] Reference Figure 4Embodiments of devices and methods for detecting children may include: in step S401, receiving images (e.g., RGB images) from one or more cameras (e.g., one or more monocular cameras) of a mobile system (e.g., associated with, on, or on the mobile system); and in step S402, calibrating the intrinsic and / or extrinsic parameters of the cameras. Embodiments may include: in step S403, transforming each point of the acquired image (from one or more cameras, where the image may come from a single camera or be a combination of multiple cameras) into points on a bird's-eye view image, extracting distorted points, and performing calibration; in step S404, implementing a lookup table that stores a previously calculated mapping from the position of each pixel in the original image to the corresponding position on the bird's-eye view image; and in step S405, generating the bird's-eye view image (e.g., fast / real-time or near real-time).
[0061] Figure 5 and Figure 6 Examples of devices and methods for detecting children are shown.
[0062] Reference Figure 5 Embodiments of the device and method for detecting children may include: in step S501, providing the first foot pixel coordinates p foot In step S502, a homography matrix for transforming the bird's-eye view image is received; and in step S503, the coordinates p of the first foot pixel in the bird's-eye view image are obtained. foot The corresponding second pixel coordinate p' foot An embodiment may include: in step S504 (e.g., after step S501), receiving an algorithm for estimating distance based on a bird's-eye view image; and in step S505, estimating the distance d between the mobile system and a person detected in the image. person .
[0063] Reference Figure 6 Step S505 may include: estimating the distance d using the following equation 1. person .
[0064] (Equation 1)
[0065]
[0066] Here, d can represent distance d person C'P' can be calculated using Equation 2 below.
[0067] (Equation 2)
[0068]
[0069] Here, CC' can represent the height of the camera, θtilt The tilt angle of the camera can be represented by v, and the first pixel coordinate p can be represented by v. foot The y-coordinate of PP' can be calculated using Equation 3 below.
[0070] (Equation 3)
[0071]
[0072] Here, u can represent the first pixel coordinate p. foot The x-coordinate, CP' can be calculated using Equation 4 below, and Cp' can be calculated using Equation 5 below.
[0073] (Equation 4)
[0074]
[0075] (Equation 5)
[0076]
[0077] Here, CC' can represent the height of the camera, and v can represent the first pixel coordinate p. foot The y-coordinate, C'P', can be calculated using Equation 2 above.
[0078] Figure 7 Examples of devices and methods for detecting children are shown.
[0079] Reference Figure 7 Embodiments of the device and method for detecting children may include: in step S701, receiving head pixel coordinates p head In step S702, the distance d between the mobile system and the person detected in the image is received. person In step S703, a pixel-to-world coordinate system transformation is performed; and in step S704, based on distance d... person and head pixel coordinates p head Estimate the height h of the person detected in the image. person .
[0080] Step S704 may include: obtaining the head pixel coordinates p of the image using Equation 6 below. head The Z-estimation of the height h is obtained by transforming (x, y) to the world coordinate system and obtaining coordinates (X, Y, Z). person .
[0081] (Equation 6)
[0082]
[0083] Here, K can represent the intrinsic parameter matrix of the camera, [R|t] can represent the extrinsic parameter matrix of the camera, and s can represent the distance d. person .
[0084] Figure 8 Examples of devices and methods for detecting children are shown.
[0085] Reference Figure 8 Embodiments of the device and method for detecting children may include: in step S801, receiving data based on distance d. person and head pixel coordinates p head Estimated height h of the person detected in the image person ; and in step S802, the height h person Compared with the predetermined reference height h kid (For example, 120cm) for comparison. An embodiment may include: in step S803, when height h... person Less than or equal to the reference height h kid At that time, it is determined that the person detected in the image is a child; and in step S804, when the height h person Exceeding the reference height h kid At that time, it was determined that the person detected in the image was an adult.
[0086] Figure 9 Examples of devices and methods for detecting children are shown.
[0087] Reference Figure 9 An embodiment of the device and method for detecting children may include: in step S901, receiving the result of whether the recipient is a child; and in step S902, receiving the second foot pixel coordinate p' of the bird's-eye view image. foot In step S903, a bird's-eye view image is received; in step S904, a bird's-eye view monitoring image is generated; and in step S905, the bird's-eye view monitoring image is output. The output of the bird's-eye view monitoring image may be based on determining that the detected person is a child. Additionally or optionally, the method disclosed herein may include controlling the operation of the mobile system (e.g., speed, movement, etc.) based on rules / safety policies associated with detecting a child, as opposed to detecting an adult. For example, the rules / safety policies may include causing the mobile system to move slower, avoid the detected child by a greater distance, and stop for a longer period at a greater distance from pedestrian-marked areas (e.g., sidewalks, crosswalks, bike lanes, etc.) based on the determination in S901 that the detected person is a child. The device 10 may receive the rules / safety policies (e.g., via a network, via user input), and the rules / safety policies may be configured to cause different operations of the mobile system depending on whether the detected person is determined to be an adult or a child.
[0088] Figure 10 This is a diagram illustrating a computing device according to an example.
[0089] Reference Figure 10 The method and apparatus for detecting children according to the example can be implemented by utilizing computing device 50 (e.g., as described herein, device 10 may include computing device 50).
[0090] The computing device 50 may include at least one processor 510, a memory 530, a user interface input device 540, a user interface output device 550, and / or a storage device 560. The components of the computing device 50 may be configured to communicate with each other via a bus 520. The computing device 50 may additionally or optionally include a network interface 570 communicatively connected to a network 40 (e.g., an external network such as the Internet, vehicle-to-vehicle networks, vehicle-to-infrastructure networks, or vehicle-to-everything networks). The network interface 570 can send signals to or receive signals from other entities via the network 40.
[0091] At least one processor 510 can be implemented as one or more of a variety of types, such as a microcontroller unit (MCU), application processor (AP), central processing unit (CPU), graphics processing unit (GPU), neural processing unit (NPU), and / or quantum processing unit (QPU). At least one processor 510 can be any (e.g., semiconductor) device configured to execute instructions / commands stored in memory 530 or storage device 560. At least one processor 510 can be configured to implement the above-mentioned... Figures 1 to 9 Describe the functions and / or methods.
[0092] The memory 530 and storage device 560 may include various types of volatile or non-volatile storage media. For example, the memory 530 may include read-only memory (ROM) 531 and random access memory (RAM) 532. In some examples, the memory 530 may be located inside or outside the processor 510, and the memory 530 may be communicatively connected to at least one processor 510.
[0093] In some examples, at least some components or functions of the child detection method and apparatus according to the examples can be implemented by instructions (e.g., programs and / or software) executed by computing device 50, and the instructions (e.g., programs and / or software) can be stored in a computer-readable medium. According to the examples, a non-transitory computer-readable medium can record instructions (e.g., programs and / or software) for performing steps included in the implementation of the child detection method and apparatus according to the examples on a computer including processor 510, which is configured to execute instructions (e.g., programs and / or commands) stored in memory 530 and / or storage device 560.
[0094] In some examples, at least some components or functions of the methods and apparatus for detecting children according to the examples may be implemented by utilizing the hardware or circuitry of computing device 50, or may be implemented as separate hardware or circuitry that can be electrically connected to computing device 50. For example, apparatus 10 may include computing device 50, which may include one or more processors and a memory storing instructions (e.g., programs and / or commands) that, when executed by one or more processors, cause computing device 50 to perform the methods disclosed herein.
[0095] This disclosure attempts to provide a method and apparatus for detecting children, which is capable of detecting adults and children separately and avoiding collisions in a mobile system that enables mobility.
[0096] According to an example, a method for detecting a child in a mobile system moving in a space where a person is present includes: acquiring an image from a camera set in the mobile system; extracting first foot pixel coordinates corresponding to the person's feet and head pixel coordinates corresponding to the person's head from the acquired image by utilizing semantic segmentation; generating a bird's-eye view image from the acquired image; acquiring second foot pixel coordinates corresponding to the first foot pixel coordinates from the bird's-eye view image; estimating the distance between the mobile system and the detected person based on the first foot pixel coordinates; and estimating the height of the detected person based on the distance and the head pixel coordinates.
[0097] In some examples, the method may further include: determining that the detected person is a child when the height value is within a predetermined range; and determining that the detected person is not a child when the height value is outside the range.
[0098] In some examples, the method may further include: displaying the second foot pixel coordinates on the bird's-eye view image and determining whether the detected person is a child.
[0099] In some examples, extracting the first foot pixel coordinates and the head pixel coordinates may include: setting the region corresponding to the human foot detected from the acquired image as a first temporary region; and extracting the coordinates of the bottommost pixel among one or more pixels included in the first temporary region as the first foot pixel coordinates.
[0100] In some examples, extracting the first foot pixel coordinates and the head pixel coordinates may include: setting the region corresponding to the head of the person detected from the acquired image as a second temporary region; and extracting the coordinates of the topmost pixel among one or more pixels included in the second temporary region as the head pixel coordinates.
[0101] In some examples, generating a bird's-eye view image may include: obtaining the intrinsic and extrinsic parameters of the camera; transforming each point on the acquired image into a point on the bird's-eye view image using a homography matrix; generating a lookup table based on the transformation result; and generating the bird's-eye view image from the acquired image in real time by utilizing the lookup table.
[0102] In some examples, obtaining the second-foot pixel coordinates may include obtaining the second-foot pixel coordinates corresponding to the first-foot pixel coordinates through a homography matrix.
[0103] In some examples, estimating distance may include: setting the first foot pixel coordinates and the camera coordinates in a normalized coordinate system; setting coordinates in a world coordinate system corresponding to the coordinates set in the normalized coordinate system; and estimating the distance between the moving system and the detected person based on the coordinates set in the world coordinate system corresponding to the first foot pixel coordinates and the camera coordinates.
[0104] In some examples, estimating the distance can include estimating the distance using Equation 1 below:
[0105] (Equation 1)
[0106]
[0107] Here, d represents the distance, and C'P' is calculated using Equation 2 below:
[0108] (Equation 2)
[0109]
[0110] Here, CC' represents the height of the camera, θ tilt Let represent the camera's tilt angle, v represent the y-coordinate of the first pixel, and PP' be calculated using Equation 3 below:
[0111] (Equation 3)
[0112]
[0113] Here, u represents the x-coordinate of the first pixel coordinate, CP' is calculated using Equation 4 below, and Cp' is calculated using Equation 5 below:
[0114] (Equation 4)
[0115]
[0116] (Equation 5)
[0117]
[0118] Here, CC' represents the height of the camera, v represents the y-coordinate of the first pixel coordinate, and C'P' is calculated using Equation 2 above.
[0119] In some examples, estimating height may include estimating the Z-coordinate of the coordinates (X, Y, Z) obtained by transforming the head pixel coordinates (x, y) to the world coordinate system using Equation 6 below as an example:
[0120] (Equation 6)
[0121]
[0122] Here, K represents the intrinsic parameter matrix of the camera, [R|t] represents the extrinsic parameter matrix of the camera, and s represents the distance.
[0123] According to an example, a device is provided for detecting a child in a mobile system, the mobile system executing program code loaded on one or more memory devices via one or more processors and moving in a space where a person is present, wherein the program code can be executed to: acquire images from a camera set in the mobile system; extract first foot pixel coordinates corresponding to the person's feet and head pixel coordinates corresponding to the person's head from the acquired images by utilizing semantic segmentation; generate a bird's-eye view image from the acquired images; acquire second foot pixel coordinates corresponding to the first foot pixel coordinates from the bird's-eye view image; estimate the distance between the mobile system and the detected person based on the first foot pixel coordinates; and estimate the height of the detected person based on the distance and the head pixel coordinates.
[0124] In some examples, the program code can be executed to: determine that the detected person is a child when the height value is within a predetermined range; and determine that the detected person is not a child when the height value is outside the range.
[0125] In some examples, the program code can be executed to: display the second foot pixel coordinates on a bird's-eye view image and determine whether the detected person is a child.
[0126] In some examples, extracting the first foot pixel coordinates and the head pixel coordinates may include: setting the region corresponding to the human foot detected from the acquired image as a first temporary region; and extracting the coordinates of the bottommost pixel among one or more pixels included in the first temporary region as the first foot pixel coordinates.
[0127] In some examples, extracting the first foot pixel coordinates and the head pixel coordinates may include: setting the region corresponding to the head of the person detected from the acquired image as a second temporary region; and extracting the coordinates of the topmost pixel among one or more pixels included in the second temporary region as the head pixel coordinates.
[0128] In some examples, generating a bird's-eye view image may include: obtaining the intrinsic and extrinsic parameters of the camera; transforming each point on the acquired image into a point on the bird's-eye view image using a homography matrix; generating a lookup table based on the transformation result; and generating the bird's-eye view image from the acquired image in real time by utilizing the lookup table.
[0129] In some examples, obtaining the second-foot pixel coordinates may include obtaining the second-foot pixel coordinates corresponding to the first-foot pixel coordinates through a homography matrix.
[0130] In some examples, estimating distance may include: setting the first foot pixel coordinates and the camera coordinates in a normalized coordinate system; setting coordinates in a world coordinate system corresponding to the coordinates set in the normalized coordinate system; and estimating the distance between the moving system and the detected person based on the coordinates set in the world coordinate system corresponding to the first foot pixel coordinates and the camera coordinates.
[0131] In some examples, estimating the distance can include estimating the distance using Equation 1 below:
[0132] (Equation 1)
[0133]
[0134] Here, d represents the distance, and C'P' is calculated using Equation 2 below:
[0135] (Equation 2)
[0136]
[0137] Here, CC' represents the height of the camera, θ tilt Let represent the camera's tilt angle, v represent the y-coordinate of the first pixel, and PP' be calculated using Equation 3 below:
[0138] (Equation 3)
[0139]
[0140] Here, u represents the x-coordinate of the first pixel coordinate, CP' is calculated using Equation 4 below, and Cp' is calculated using Equation 5 below:
[0141] (Equation 4)
[0142]
[0143] (Equation 5)
[0144]
[0145] Here, CC' represents the height of the camera, v represents the y-coordinate of the first pixel coordinate, and C'P' is calculated using Equation 2 above.
[0146] In some examples, estimating height may include estimating the Z-coordinate of the coordinates (X, Y, Z) obtained by transforming the head pixel coordinates (x, y) to the world coordinate system using Equation 6 below as an example:
[0147] (Equation 6)
[0148]
[0149] Here, K represents the intrinsic parameter matrix of the camera, [R|t] represents the extrinsic parameter matrix of the camera, and s represents the distance.
[0150] Based on the information disclosed herein, it is possible to accurately measure distances to people using only a monocular camera without employing expensive equipment such as LiDAR or RGB-D cameras, and to prevent collisions with people during autonomous driving or traffic control operations in mobile systems. Specifically, by detecting people in bird's-eye view images and distinguishing between adults and children, collision avoidance strategies for mobile systems can be refined accordingly, and customized services and content related to mobility can be provided.
[0151] Although examples of this disclosure have been described in detail above, the scope of this disclosure is not limited thereto, and various modifications and improvements made by those skilled in the art to which this disclosure pertains also fall within the scope of this disclosure.
Claims
1. A method for controlling the operation of a vehicle, the method comprising: Images are acquired from the vehicle's camera; Based on the semantic segmentation of the image, extract: The first foot pixel coordinates corresponding to the human foot detected in the acquired image; as well as Head pixel coordinates corresponding to the person's head; A transformed view image is generated by converting the coordinate system of the acquired image to the coordinate system of the transformed view. Based on the transformed view image, determine the coordinates of the second foot pixel in the transformed view coordinate system that correspond to the coordinates of the first foot pixel; The distance between the vehicle and the detected person is estimated based on the first foot pixel coordinates; the height of the detected person is estimated based on the distance and the head pixel coordinates; as well as The vehicle is controlled based on the estimated height.
2. The method according to claim 1, further comprising: Based on the height value being within a predetermined range, it is determined that the detected person is a child; or Based on the value being outside the predetermined range, it is determined that the detected person is not a child.
3. The method according to claim 2, further comprising: Displayed on the transformed view image: The second pixel coordinate; as well as The result determines whether the detected person is a child.
4. The method according to claim 1, wherein, Extracting the first foot pixel coordinates includes: The first temporary region is set as the region corresponding to the human feet detected in the acquired image; and Extract the coordinates of the bottommost pixel from one or more pixels in the first temporary region as the coordinates of the first foot pixel.
5. The method according to claim 1, wherein, Extracting the head pixel coordinates includes: The second temporary region is set as the region corresponding to the head of a person detected in the acquired image; and Extract the coordinates of the topmost pixel from one or more pixels in the second temporary region as the head pixel coordinates.
6. The method according to claim 1, wherein, Generating the transformed view image includes: Obtain the intrinsic and extrinsic parameters of the camera; By using the homography matrix based on the intrinsic and extrinsic parameters, each point in the acquired image is transformed into a point in the transformed view image; A lookup table is generated based on the transformation; and The transformed view image is generated based on the lookup table and the acquired image.
7. The method of claim 6, further comprising: Based on the homography matrix, the coordinates of the second foot pixel corresponding to the coordinates of the first foot pixel are obtained.
8. The method according to claim 1, wherein, Estimating the distance includes: Set the coordinates of the first foot pixel and the coordinates of the camera in a normalized coordinate system; Determine the corresponding coordinates in the world coordinate system, wherein the corresponding coordinates correspond to the first foot pixel coordinates and the camera coordinates in the normalized coordinate system; and Based on the corresponding coordinates in the world coordinate system, the distance between the vehicle and the detected person is estimated.
9. The method according to claim 1, wherein, Estimating the distance includes: The distance is estimated using Equation 1 below: Equation 1 d represents the distance, and C'P' is calculated using Equation 2: Equation 2 CC' represents the height of the camera, θ tilt The tilt angle of the camera is represented by v, which represents the y-coordinate of the first pixel. PP' is calculated using Equation 3: Equation 3 u represents the x-coordinate of the first foot pixel coordinate. CP' is calculated using Equation 4, and Cp' is calculated using Equation 5. Equation 4 Equation 5 CC' represents the height of the camera, v represents the y-coordinate of the first foot pixel coordinate, and C'P' is calculated using Equation 2.
10. The method according to claim 1, wherein, The estimated height includes: The height is estimated by the value of Z in the coordinates (X, Y, Z) obtained by transforming the head pixel coordinates (x, y) to the world coordinate system using Equation 6. Equation 6 K represents the intrinsic parameter matrix of the camera, [R|t] represents the extrinsic parameter matrix of the camera, and s represents the distance.
11. A vehicle apparatus, the apparatus comprising one or more processors, the one or more processors executing program code loaded on one or more memory devices, wherein, The program code configures the device when executed by the one or more processors: Images are acquired from the vehicle's camera; Based on the semantic segmentation of the image, extract: The first foot pixel coordinates corresponding to the human foot detected in the image; as well as Head pixel coordinates corresponding to the person's head; A transformed view image is generated by converting the coordinate system of the acquired image to the coordinate system of the transformed view. Based on the transformed view image, determine the coordinates of the second foot pixel in the transformed view coordinate system that correspond to the coordinates of the first foot pixel; The distance between the vehicle and the detected person is estimated based on the coordinates of the first foot pixel; The height of the detected person is estimated based on the distance and the head pixel coordinates; and The vehicle is controlled based on the estimated height.
12. The apparatus according to claim 11, wherein, The program code configures the device when executed by the one or more processors: Based on the height value being within a predetermined range, it is determined that the detected person is a child; or Based on the value being outside the predetermined range, it is determined that the detected person is not a child.
13. The apparatus according to claim 12, wherein, The program code configures the device when executed by the one or more processors: Displayed on the transformed view image: The second foot pixel coordinates; and The result determines whether the detected person is a child.
14. The apparatus according to claim 11, wherein, The program code configures the device when executed by the one or more processors: The first foot pixel coordinates are extracted as follows: The first temporary region is set as the region corresponding to the human feet detected in the acquired image; and Extract the coordinates of the bottommost pixel from one or more pixels in the first temporary region as the coordinates of the first foot pixel.
15. The apparatus according to claim 11, wherein, The program code configures the device when executed by the one or more processors: The head pixel coordinates are extracted as follows: The second temporary region is set as the region corresponding to the head of a person detected in the acquired image; and Extract the coordinates of the topmost pixel from one or more pixels in the second temporary region as the head pixel coordinates.
16. The apparatus according to claim 11, wherein, The program code configures the device when executed by the one or more processors: The transformed view image is generated as follows: Obtain the intrinsic and extrinsic parameters of the camera; By using the homography matrix based on the intrinsic and extrinsic parameters, each point in the acquired image is transformed into a point in the transformed view image; A lookup table is generated based on the transformation; and The transformed view image is generated based on the lookup table and the acquired image.
17. The apparatus according to claim 16, wherein, The program code configures the device when executed by the one or more processors: Based on the homography matrix, the coordinates of the second foot pixel corresponding to the coordinates of the first foot pixel are obtained.
18. The apparatus according to claim 11, wherein, The program code configures the device when executed by the one or more processors: The distance is estimated using the following method: Set the coordinates of the first foot pixel and the coordinates of the camera in a normalized coordinate system; Determine the corresponding coordinates in the world coordinate system, wherein the corresponding coordinates correspond to the coordinates of the first foot pixel and the coordinates of the camera in the normalized coordinate system; and Based on the corresponding coordinates in the world coordinate system, the distance between the vehicle and the detected person is estimated.
19. The apparatus according to claim 11, wherein, Estimating the distance includes: The distance is estimated using Equation 1 below: Equation 1 d represents the distance, and C'P' is calculated using Equation 2: Equation 2 CC' represents the height of the camera, θ tilt The tilt angle of the camera is represented by v, which represents the y-coordinate of the first pixel. PP' is calculated using Equation 3: Equation 3 u represents the x-coordinate of the first foot pixel coordinate. CP' is calculated using Equation 4, and Cp' is calculated using Equation 5. Equation 4 Equation 5 CC' represents the height of the camera, v represents the y-coordinate of the first foot pixel coordinate, and C'P' is calculated using Equation 2.
20. The apparatus according to claim 11, wherein, The estimated height includes: The height is estimated by the value of Z in the coordinates (X, Y, Z) obtained by transforming the head pixel coordinates (x, y) to the world coordinate system using Equation 6. Equation 6 K represents the intrinsic parameter matrix of the camera, [R|t] represents the extrinsic parameter matrix of the camera, and s represents the distance.
Citation Information
Patent Citations
ICT / IoT based Aeration Device for the Treatment of Livestock Manure Control System
KR1020240061714A