Object Recognition Method Of Markerless-Based Gait Image And Computer Readable Recording Medium Recording A Program Performing The Same

The markerless-based object recognition method preprocesses walking images using semantic segmentation and shooting guides to address frame imbalance and background interference, enhancing AI accuracy for musculoskeletal analysis.

KR102994143B1Active Publication Date: 2026-07-21INDUSTRYACADEMIC COOPERATION FOUNDATION GYEONGSANG NATIONAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
INDUSTRYACADEMIC COOPERATION FOUNDATION GYEONGSANG NATIONAL UNIVERSITY
Filing Date
2023-03-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing AI-based object recognition technologies for musculoskeletal analysis from gait images struggle with accuracy due to issues like frame imbalance, background interference, and incomplete capture of the pedestrian's body, especially in non-professional shooting environments without markers.

Method used

A markerless-based object recognition method that preprocesses walking images using semantic segmentation and provides shooting guides to ensure the entire body is captured, employing techniques like DeeplabV3, ResNet-101, YOLO, and coordinate averaging to improve accuracy.

Benefits of technology

Enhances the accuracy of AI-based object recognition by minimizing errors and ensuring complete body capture, thereby improving musculoskeletal disorder diagnosis in everyday settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112023032477417-PAT00001_ABST
    Figure 112023032477417-PAT00001_ABST
Patent Text Reader

Abstract

The present invention relates to a markerless-based object recognition method for walking images and a computer-readable recording medium having a program for performing the same. The method comprises: a walking image acquisition step in which a walking image of a pedestrian walking without a marker is acquired; a preprocessing step in which the walking image is preprocessed; and an object recognition step in which a pedestrian is recognized as an object for each frame of the preprocessed walking image. The invention also relates to a computer-readable recording medium having a program for performing the same.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a markerless-based object recognition method for walking images and a computer-readable recording medium having a program for performing the same. More specifically, it relates to a markerless-based object recognition method for walking images that preprocesses walking images to improve the accuracy of artificial intelligence-based object recognition technology, and a computer-readable recording medium having a program for performing the same. Background Technology

[0002] Recently, artificial intelligence technology is being utilized in various fields, including autonomous driving, aviation, and agriculture. In particular, the medical sector is actively developing AI technologies based on accumulated big data to assist doctors in their decision-making. Such AI is expected to become an effective treatment method, as it enables immediate basic diagnosis and urgency analysis of patients at a low cost.

[0003] Among them, orthopedics utilizes gait images and videos to analyze gait characteristics useful for determining musculoskeletal diseases and their severity. In this regard, Related Literature 1 describes a device and application for predicting musculoskeletal abnormalities, which is a technology capable of determining whether a gait is abnormal by acquiring a 2D gait video and inputting it into a machine learning model. However, it may be difficult to accurately recognize the pedestrian as an object in gait videos depending on the pedestrian's shooting environment and shooting method.

[0004] Related Literature 2 relates to a method and device for acquiring motion for musculoskeletal diagnosis, which can acquire motion for musculoskeletal diagnosis based on wearable multi-sensors and depth images using pre-measured human body information. However, since pedestrians must wear wearable multi-sensors to acquire images, while this system can be used to diagnose the musculoskeletal system when visiting specific institutions such as orthopedic clinics, it may be difficult for individuals to receive a musculoskeletal diagnosis on a personal basis in their daily lives.

[0005] Therefore, there is an urgent need in this field for technology capable of accurately determining musculoskeletal disorders and their severity even in general situations where no markers are attached to a pedestrian's body. Prior art literature

[0006] Korean Registered Patent Document No. 10-2251925 Korean Registered Patent Document No. 10-2020-0087027 The problem to be solved

[0007] The present invention aims to solve the aforementioned problems and to obtain a markerless-based object recognition method for walking images in which walking images are preprocessed using a semantic segmentation technique to improve the accuracy of artificial intelligence-based object recognition technology, and a computer-readable recording medium having a program for performing the same.

[0008] In addition, the objective of the present invention is to provide a markerless-based object recognition method for walking images in which a shooting guide for walking images is provided so that the entire body of a pedestrian is included in each frame of the walking image to prevent object recognition errors, and a computer-readable recording medium having a program for performing the same recorded thereon.

[0009] The technical problems that the present invention aims to solve are not limited to those mentioned above, and other unmentioned technical problems can be clearly understood by those skilled in the art from the description of the present invention. means of solving the problem

[0010] To achieve the above objective, the object recognition method for markerless-based walking images according to the present invention comprises: a walking image acquisition step in which a walking image of a pedestrian walking without a marker is acquired by at least one processor; a preprocessing step in which the walking image is preprocessed by the at least one processor; and an object recognition step in which a pedestrian is recognized as an object for each frame of the preprocessed walking image by the at least one processor.

[0011] To achieve the above objective, the present invention relates to a computer-readable recording medium having a program recorded thereon for performing a markerless-based object recognition method for walking images. Effects of the invention

[0012] As described above, according to the present invention, by utilizing a semantic segmentation technique to preprocess walking images, there is an effect of improving the accuracy of artificial intelligence-based object recognition technology.

[0013] In addition, the present invention provides a shooting guide for walking images, thereby having the effect of preventing object recognition errors in advance by ensuring that the pedestrian's entire body is included in each frame of the walking image when shooting walking images.

[0014] The effects of the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the detailed description and claims. Brief explanation of the drawing

[0015] FIG. 1 is a flowchart of a markerless-based object recognition method for walking images according to an embodiment of the present invention. Figure 2 is a drawing showing object recognition errors caused by frame imbalance (a), clothing (b), and background (c). FIG. 3 is a drawing showing a frame (a) in which an object recognition error occurred due to the background and a frame (b) that has been background-processed from a preprocessing step according to an embodiment of the present invention. FIG. 4 is a drawing showing the case where the body part above the pedestrian's neck is outside the frame (a) and the case where the arm or leg is outside the frame when the pedestrian walks to the right (b). FIG. 5 is a drawing showing a frame (a) with a bounding box displayed from a bounding box derivation step according to an embodiment of the present invention, and coordinate values ​​for a first vertex and a second vertex derived from a coordinate value derivation step. FIG. 6 is a drawing showing an ideal shooting guide (a) and a maximum shooting guide (b) according to an embodiment of the present invention. Specific details for implementing the invention

[0016] The terms used in this specification have been selected based on currently widely used general terms whenever possible, taking into account their functions in the present invention; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this invention should be defined not merely by their names, but based on their meanings and the overall content of the invention.

[0017] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0018] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings. FIG. 1 is a flowchart of a markerless-based object recognition method for walking images according to an embodiment of the present invention. FIG. 2 is a diagram showing object recognition errors caused by frame imbalance (a), clothing (b), and background (c). FIG. 3 is a diagram showing a frame (a) where an object recognition error occurred due to the background and a frame (b) that has been background-processed from a preprocessing step (S200) according to an embodiment of the present invention. FIG. 4 is a diagram showing a case where a body part above the pedestrian's neck is outside the frame (a) and a case where an arm or leg is outside the frame when the pedestrian walks to the right (b). FIG. 5 is a diagram showing a frame (a) with a bounding box displayed from a bounding box derivation step (S420) according to an embodiment of the present invention and coordinate values ​​for a first vertex and a second vertex derived from a coordinate value derivation step (S430). FIG. 6 is a drawing showing an ideal shooting guide (a) and a maximum shooting guide (b) according to an embodiment of the present invention.

[0019] First, the present invention includes a recording medium (120) readable by a computer device (100) on which a program for performing a markerless-based object recognition method for walking images is recorded. For example, it may be a CD, DVD, hard disk, Blu-ray disk, USB, memory card, ROM, etc. And the markerless-based object recognition method for walking images of the present invention may be implemented by at least one processor (110) in the computer device (100) reading the recording medium (120).

[0020] Referring to FIG. 1, the object recognition method for markerless-based walking images of the present invention includes a walking image acquisition step (S100) in which a walking image of a pedestrian walking without a marker is acquired by at least one processor (110), a preprocessing step (S200) in which the walking image is preprocessed by the at least one processor (110), and an object recognition step (S300) in which a pedestrian is recognized as an object for each frame of the preprocessed walking image by the at least one processor (110).

[0021] The walking video referred to in the present invention is a markerless-based video acquired from a pedestrian who is not equipped with a marker. Generally, a pedestrian's walking can be tracked using a system that includes a separate marker, which is a type of sensor, on a part of the pedestrian's body, and another sensor for recognizing it. However, pedestrians face the inconvenience of having to wear multiple markers. To solve this, the walking video acquisition step (S100) may acquire a walking video in which the walking motion of the pedestrian is captured through a camera or a device with a built-in camera, such as a smartphone.

[0022] In addition, the above-mentioned walking video is characterized by being a video in which the left or right side of the pedestrian is captured. The present invention aims to improve the accuracy of an artificial intelligence model capable of predicting musculoskeletal diseases and the severity of the disease. As described above, since the walking video is a two-dimensional image acquired through a shooting device, it is difficult to accurately identify the walking motion of a pedestrian using a video in which the front or rear of the pedestrian is captured. That is, in order to determine musculoskeletal diseases more efficiently and accurately from the two-dimensional walking video, it is most preferable that the walking video acquisition step (S100) acquires a video in which the right or left side is captured, rather than a video in which the front or rear of the pedestrian is captured.

[0023] Next, the preprocessing step (S200) is intended to minimize recognition errors that may occur from the object recognition step (S300). As described above, since the pedestrian video is captured by a non-professional using a general shooting device, various problems may occur. These include a frame imbalance problem where part of the pedestrian's body is out of frame in the video, as shown in FIG. 2 (a); a clothing problem where the pedestrian's body shape cannot be accurately identified due to wearing clothes that are larger than the body, as shown in FIG. 2 (b); and a background problem where the pedestrian cannot be accurately identified because various objects other than the pedestrian are captured depending on the shooting environment, as shown in FIG. 2 (c). In particular, the present invention aims to solve the frame imbalance problem and the background problem through an information processing method.

[0024] To solve background problems, the above preprocessing step (S200) is characterized by including a frame extraction step (S210), a pedestrian detection step (S220), a binarization step (S230), an area calculation step (S240), and a background processing step (S250).

[0025] In the above frame extraction step (S210), multiple frames can be extracted from the walking video. The walking video is not an image format formed of a single frame, but a video format formed of multiple frames. That is, because the appearance of the pedestrian differs from frame to frame, the above frame extraction step (S210) separates and extracts the walking video into multiple frames in order to identify this.

[0026] Next, the pedestrian detection step (S220) utilizes a semantic segmentation technique to detect pedestrians in each frame. The semantic segmentation technique is a method for finding the boundaries between objects by dividing each frame of the pedestrian video into multiple sets of pixels. Most preferably, the pedestrian detection step (S220) utilizes a semantic segmentation technique based on a DeeplabV3 model using ResNet-101 as the backbone so that objects and the background outside the objects can be distinguished. In other words, the pedestrian detection step (S220) is a process of distinguishing objects and the background outside the objects.

[0027] Next, the binarization step (S230) may binarize the detected pedestrian. If there are no objects in the pedestrian image and only pedestrians are present, the semantic segmentation technique alone may be sufficient to detect the pedestrian. However, if there are objects in the pedestrian image due to the shooting environment, the objects distinguished by the semantic segmentation technique from the pedestrian detection step (S220) may include both pedestrians and non-pedestrian objects. Therefore, the binarization step (S230) is intended to distinguish between pedestrians and non-pedestrian objects among the objects. According to one embodiment of the present invention, the binarization step (S230) may binarize only the pedestrians among the objects into white.

[0028] Next, in the area calculation step (S240), a contour detection technique is used to designate a contour area containing a pedestrian, and the area for the contour area for each frame can be calculated. As described above, since the movement of the pedestrian varies by frame, the pedestrian's area may vary by frame. In the area calculation step (S240), a contour detection technique is used on the area binarized from the binarization step (S230) so that the contour area can be designated as a single line. Furthermore, in the area calculation step (S240), the area can be calculated on the contour area for each frame.

[0029] Next, in the background processing step (S250), the background, which is the remaining part excluding the maximum contour area in each frame, can be processed as a solid color. And the object recognition step (S300) is characterized by recognizing the maximum contour area as an object. Contour areas exist for each frame, and each contour area has a different area. The background processing step (S250) targets the maximum contour area among them. This is to provide a margin to the area so that the pedestrian's body can be entirely contained within the contour area.

[0030] Looking at Fig. 3(a), it can be seen that an object recognition error occurs in an arbitrary frame due to the background before preprocessing. Looking at Fig. 3(b), after preprocessing, the pedestrian and the background other than the pedestrian can be accurately distinguished, and the background can be binarized into black. Accordingly, it can be seen that no object recognition error occurs due to the background.

[0031] Next, the object recognition method for a markerless-based walking image of the present invention further comprises a shooting guide providing step (S400) in which a shooting guide for a walking image is provided by the at least one processor (110) so that the entire body of a pedestrian can be included in each frame of the walking image.

[0032] To resolve the frame imbalance problem, the above shooting guide providing step (S400) may include a data set acquisition step (S410), a bounding box derivation step (S420), a coordinate value derivation step (S430), a frame average value calculation step (S440), and a walking video average value calculation step (S450).

[0033] The frame imbalance mentioned in the present invention refers to a case where a part of a pedestrian's body, such as the neck or legs, extends beyond the frame (F). Looking at FIG. 4 (a), it can be seen that the boundary box (B) also extends beyond the frame (F) as the body part above the neck extends beyond the frame (F). Looking at FIG. 4 (b), it can also be seen that the boundary box (B) extends beyond the frame (F) as the arm or leg extends beyond the frame (F).

[0034] Next, the data set acquisition step (S410) may acquire multiple walking images, each containing the full body of a pedestrian in each frame, as a data set from the walking image acquisition step (S100). It is most preferable that the data set acquisition step (S410) acquires walking images that include the full body of a pedestrian in each frame, rather than walking images that include frames where parts of the pedestrian's body are out of view. This is to provide a shooting guide to any pedestrian who will subsequently have their walking images captured through a shooting device, so it is most preferable that normal walking images be the subject of the data set.

[0035] Next, the bounding box derivation step (S420) utilizes the YOLO algorithm to derive a pedestrian-centered bounding box for each frame within multiple pedestrian images. The YOLO (You Only Look Once) algorithm mentioned in the present invention has the effect of enabling rapid object classification because it can process a single frame as a whole at once without dividing and interpreting each frame into multiple pieces. The bounding box derivation step (S420) utilizes the YOLO algorithm to derive a bounding box centered on the classified pedestrian. The bounding box (B) is as shown in FIG. 5 (a).

[0036] Next, in the coordinate value derivation step (S430), the coordinate values ​​for the first vertex (P1) and the second vertex (P2) facing diagonally in the bounding box (B) for each frame (F) can be derived, respectively. In the embodiment of FIG. 5 (b), the first vertex (P1) may be the upper left vertex (x,y), and the second vertex (P2) may be the lower right vertex (x+w, y+h). That is, in the coordinate value derivation step (S430), only two vertices (P1, P2) are calculated instead of all four vertices, so that a bounding box (B) of size w*h can be identified, which has a significant effect of reducing the amount of calculation and increasing the calculation speed.

[0037] Next, the frame average value calculation step (S440) may calculate a first frame average value (AVGF1) by averaging the first vertex (P1) of the bounding box (B) for each frame and a second frame average value (AVGF2) by averaging the second vertex (P2). For example, any walking video may consist of frames F1 to F5. Then, the first frame average value (AVGF1) is the value obtained by averaging the first vertex (P1_1) of frame F1, the first vertex (P1_2) of frame F2, the first vertex (P1_3) of frame F3, the first vertex (P1_4) of frame F4, and the first vertex (P1_5) of frame F5. And the second frame average value (AVGF2) is the average of the second vertex of frame F1 (P2_1), the second vertex of frame F2 (P2_2), the second vertex of frame F3 (P2_3), the second vertex of frame F4 (P2_4), and the second vertex of frame F5 (P2_5). If there are 20 walking videos in the dataset, the operation for each walking video can be repeated as described above.

[0038] Next, the step of calculating the average value of the walking video (S450) may calculate the first average value of the walking video (AVGV1) obtained by averaging the first frame average value (AVGF1) for each walking video, and the second average value of the walking video (AVGV2) obtained by averaging the second frame average value (AVGF2). For example, if the data set contains walking videos V1 to V20, the first average value of the walking video (AVGV1) is the value obtained by averaging the first frame average value (AVGF1_1) for walking video V1 to the first frame average value (AVGF1_20) for walking video V20. And the second average value of the walking video (AVGV2) is the value obtained by averaging the second frame average value (AVGF2_1) for walking video V1 to the second frame average value (AVGF2_20) for walking video V20. That is, the size of the bounding box (B) averaged over the entire data set can be calculated from the frame average value calculation step (S440) and the walking image average value calculation step (S450).

[0039] At this time, the shooting guide provision step (S400) is characterized in that a bounding box formed by the first pedestrian image average value (AVGV1) and the second pedestrian image average value (AVGV2) is provided as an ideal shooting guide (G_Ideal). FIG. 6 (a) is the ideal shooting guide (G_Ideal), and it is most desirable for a pedestrian within the guide to be shot.

[0040] Additionally, the shooting guide providing step (S400) further includes a condition boundary box selection step (S460) in which a condition boundary box satisfying a preset condition is selected among a plurality of boundary boxes (B), and the condition boundary box formed by a first vertex having a minimum value and a second vertex having a maximum value among the first vertex and the second vertex of the condition boundary box is provided as a maximum shooting guide (G_Maximum).

[0041] There are three pre-set conditions mentioned in the present invention. The first condition is when the x-coordinate value of the first vertex (x,y) of an arbitrary bounding box (B) is smaller than the x-coordinate value of the first average value of the walking image (AVGV1). The second condition is when the x+w coordinate value of the second vertex (x+w, y+h) of an arbitrary bounding box (B) is larger than the x-coordinate value of the second average value of the walking image (AVGV2). The third condition is when the y+h coordinate value of the second vertex (x+w, y+h) of an arbitrary bounding box (B) is larger than the y-coordinate value of the second average value of the walking image (AVGV2). The three conditions described above are conditions for selecting a bounding box that is larger than the bounding box (B) obtained by averaging the data set among a plurality of bounding boxes (B).

[0042] The condition bounding box is characterized by being provided as a Maximum shooting guide (G_Maximum) formed by the first vertex (P1) having the minimum value and the second vertex (P2) having the maximum value among the first vertex (P1) and the second vertex (P2) of the condition bounding box selected through the three conditions described above. The Maximum shooting guide (G_Maximum) is an area where shooting within the frame is permitted. FIG. 6(b) is the Maximum shooting guide (G_Maximum), and it is most desirable for a pedestrian within the guide to be shot.

[0043] Therefore, it is most desirable to photograph a pedestrian when the pedestrian is positioned within the Ideal shooting guide (G_Ideal) using a general shooting device; however, if this is not possible, it may be permitted to photograph the pedestrian when the pedestrian is positioned within the Maximum shooting guide (G_Maximum). According to the present invention, by providing a shooting guide, there is an effect of preventing object recognition errors from occurring in advance.

[0044] The embodiments may be implemented by hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. Where implemented by software, firmware, middleware, or microcode, program code or code segments that perform the necessary tasks may be stored on a computer-readable storage medium and executed by one or more processors.

[0045] Furthermore, aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules or components executed by a computer. Generally, program modules or components include routines, programs, objects, and data structures that perform specific tasks or implement specific data types. The aspects of the subject matter described herein may be implemented in distributed computing environments where tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may be located on both local and remote computer storage media, including memory storage devices.

[0046] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or the components of the system, structure, device, circuit, etc. described are combined or assembled in a form different from the described method, or are replaced or substituted by other components or equivalents.

[0047] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below. Explanation of the symbols

[0048] 100.. Computer device 110.. at least one processor 120.. Recording media S100.. Walking video acquisition stage S200.. Preprocessing step S210.. Frame extraction stage S220.. Pedestrian detection stage S230.. Binarization step S240.. Area calculation step S250.. Background processing stage S300.. Object recognition stage S400.. Shooting guide provision stage S410.. Data set acquisition stage S420.. Bounding box derivation step S430.. Coordinate value derivation step S440.. Frame average value calculation step S450.. Walking video average value calculation step S460.. Condition Bounding Box Selection Step F.. Frame B.. Boundary box G_Ideal.. Ideal Shooting Guide G_Maximum.. Maximum Shooting Guide P1.. 1st vertex P2.. 2nd vertex AVGF1.. Average value of the first frame AVGF2.. Average value of the second frame AVGV1.. Average value of the first gait video AVGV2.. Average value of the second gait video

Claims

Claim 1 A shooting guide providing step in which a shooting guide is provided by at least one processor so that the entire body of a pedestrian can be included in each frame of a walking video; a walking video acquisition step in which the walking video in which a pedestrian without a marker is walking is acquired by the at least one processor; and a preprocessing step in which the walking video is preprocessed by the at least one processor. The method comprises an object recognition step in which a pedestrian is recognized as an object for each frame of a preprocessed walking video by the at least one processor; and the shooting guide providing step comprises: a data set acquisition step in which a plurality of walking videos, each containing the full body of a pedestrian in each frame, are acquired as a data set; a bounding box derivation step in which a pedestrian-centered bounding box is derived for each frame within the plurality of walking videos by using the YOLO algorithm; a coordinate value derivation step in which coordinate values ​​((x,y), (x+w, y+h)) for a first vertex (P1) and a second vertex (P2) facing diagonally in the bounding box for each frame are respectively derived; and a frame average value calculation step in which a first frame average value (AVGF1) obtained by averaging the first vertex (P1) of the bounding box for each frame and a second frame average value (AVGF2) obtained by averaging the second vertex (P2) are calculated. A method for object recognition of a markerless-based walking video, comprising: a step for calculating a walking video average value, wherein a first walking video average value (AVGV1) obtained by averaging a first frame average value (AVGF1) for each walking video and a second walking video average value (AVGV2) obtained by averaging a second frame average value (AVGF2) are calculated; wherein a bounding box formed by the first walking video average value (AVGV1) and the second walking video average value (AVGV2) is provided as an ideal shooting guide, thereby resolving the frame imbalance problem in which a part of a pedestrian's body is filmed outside the frame and preventing an object recognition error from occurring in the object recognition step. Claim 2 The method for object recognition of a markerless-based walking image according to claim 1, wherein the preprocessing step comprises: a frame extraction step in which a plurality of frames are extracted from the walking image; a pedestrian detection step in which a pedestrian is detected in each frame using a semantic segmentation technique; and a binarization step in which the detected pedestrian is binarized. Claim 3 The object recognition method for markerless-based walking images according to claim 2, further comprising: a preprocessing step in which a contour area containing a pedestrian is designated using a contour detection technique and an area for the contour area per frame is calculated; and a background processing step in which the background, which is the remaining part excluding the maximum contour area in each frame, is processed as a solid color. Claim 4 In paragraph 3, the object recognition step is characterized by recognizing the maximum contour area as an object. A markerless-based object recognition method for walking images. Claim 5 A markerless-based object recognition method for walking images, characterized in that, in claim 1, the walking image is an image in which the left or right side of the pedestrian is captured. Claim 6 The object recognition method for markerless-based walking images according to claim 1, wherein the shooting guide providing step further includes a conditional boundary box selection step in which a conditional boundary box satisfying a preset condition among a plurality of boundary boxes is selected, and a conditional boundary box formed by a first vertex having a minimum value and a second vertex having a maximum value among the first vertex and the second vertex of the conditional boundary box is provided as a maximum shooting guide. Claim 7 A computer-readable recording medium having a program for performing a markerless-based object recognition method of any one of paragraphs 1 to 6. Claim 8 delete Claim 9 delete Claim 10 delete