Method and apparatus for tracking moving object in real time by using cradle head servo motor

The method and device using a cradle head servo motor and deep learning models in an image recognition device address the challenges of tracking objects of varying sizes and rapid movements by providing two optimized modes for accuracy and speed, enhancing tracking consistency and reducing errors.

WO2025135484A1PCT designated stage expired Publication Date: 2025-06-26SOONCHUNYANG UNIV IND ACAD COOP FOUND
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017271
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-11-05
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current real-time moving object tracking systems face challenges when small and large objects coexist, leading to prioritization of larger objects and difficulties in consistent tracking across scales. Additionally, rapid object movement complicates accurate tracking, potentially resulting in errors and misidentifications.

Method used

A method and device using a cradle head servo motor in an image recognition device that employs deep learning models for object recognition and tracking. The system includes an accuracy-first mode for balancing high accuracy and speed in smartphone environments and a speed-first mode for optimizing inference speed for fast-moving objects.

Benefits of technology

The proposed solution effectively addresses the challenges of tracking objects of varying sizes and rapid movements by providing two modes that optimize accuracy and speed, thereby reducing potential errors and improving object tracking consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017271_26062025_PF_FP_ABST
    Figure KR2024017271_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method performed by an image recognition device and, more specifically, to a method and an apparatus for tracking a moving object in real time by using a cradle head servo motor, the method comprising the steps of: loading image data; preprocessing the image data; training at least one deep learning model by using the preprocessed image data; setting a mode of recognizing an object; recognizing an object by utilizing a deep learning model corresponding to the object recognition mode in accordance with the object recognition mode; and tracking the recognized object and assigning an object ID.
Need to check novelty before this filing date? Find Prior Art

Description

Real-time moving object tracking method and device using cradle head servo motor

[0001] The present invention relates to a method and device for tracking a moving object in real time using a cradle head servo motor, and more specifically, to a method and device for recognizing an object and tracking it in real time through at least one object recognition mode by using a cradle head servo motor installed in an image recognition device.

[0002] Due to the continuous improvement of smartphone performance and the spread of deep learning technology, smartphone-based object detection technology using cradle servo motors is attracting attention from researchers.

[0003] Real-time moving object detection technology is widely used in diverse fields, including robotics, transportation, manufacturing, and security. In particular, the overall market size is expected to reach $480 billion by 2027, driven by the continued growth of the creator economy. As the number of YouTube creators increases and the market grows, demand for cutting-edge technologies supporting content creation is also increasing. YouTube creators primarily rely on AI smartphone mounts to create content.

[0004] Real-time mobile object detection technology has rapidly developed due to its inherent portability and user-friendliness. However, as the speed of real-time moving objects or objects increases, the accuracy of their detection also decreases. This issue has led to numerous discussions on how to address this issue.

[0005] Furthermore, there is a growing body of research discussing ways to detect moving objects or objects differently depending on the situation, focusing on accuracy or detecting fast objects more quickly. Therefore, practical research and discussion are needed on object detection methods tailored to the characteristics of these moving objects or objects and the specific circumstances.

[0006] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired during the process of deriving the present invention, and cannot necessarily be said to be technology known to the general public prior to the filing of the present invention.

[0007] The problem to be solved by the disclosure of the present invention is that the current technical problem of real-time moving object tracking systems occurs when small and large objects coexist, which causes the algorithm to prioritize the larger object or have difficulty in tracking it consistently across various scales, and also that rapid object movement further complicates accurate tracking, which can lead to potential errors and misidentifications, and the present invention relates to a method for solving these problems.

[0008] In addition, the problem to be solved through the disclosure of the present invention relates to a method that can provide two approaches for object detection: an accuracy-first mode that can achieve a balance between high accuracy and speed required in a smartphone environment, and a speed-first mode that can emphasize optimization of inference speed for tracking fast-moving objects.

[0009] A method performed by an image recognition device for a problem to be solved through some embodiments of the present invention may include a step of loading image data, a step of preprocessing the image data, a step of learning at least one deep learning model using the preprocessed image data, a step of setting an object recognition mode, a step of recognizing an object using a deep learning model corresponding to the object recognition mode according to the object recognition mode, and a step of tracking the recognized object and assigning an object ID.

[0010] In one embodiment, the mode for recognizing the object may be characterized by including an accuracy priority mode and a speed priority mode.

[0011] In one embodiment, the deep learning model corresponding to the accuracy priority mode may be built using a CSPNet structure integrated with ResNet based on the YOLOv8 architecture.

[0012] In one embodiment, the deep learning model corresponding to the speed priority mode may be YOLOv5n built using the shallowest architecture.

[0013] In one embodiment, the step of tracking a recognized object and assigning an object ID may be characterized by assigning the same object ID when the feature map similarity between the object features in the current frame and the object features in the previous frame exceeds 50% based on the feature map.

[0014] In one embodiment, the step of recognizing an object using a deep learning model corresponding to an object recognition mode according to an object recognition mode is characterized in that the x-axis coordinate position of the object is calculated using the mathematical expression 1 in a cradle included in an image recognition device.

[0015] [Mathematical Formula 1]

[0016]

[0017] The above MD may be the distance from the center of the screen to the object, and OP may be the x-axis coordinate of the object.

[0018] In one embodiment, the step of recognizing an object using a deep learning model corresponding to the object recognition mode according to the object recognition mode is characterized in that the time required for the cradle to rotate 360 ​​degrees is calculated based on the MD value using mathematical expression 2.

[0019] [Equation 2]

[0020]

[0021] The above RT may be the time required for the cradle to rotate 360 ​​degrees.

[0022] According to the problem solving means of the present invention described above, the current technical problem of real-time moving object tracking system occurs when small objects and large objects coexist, which causes the algorithm to prioritize larger objects or have difficulty in tracking them consistently across various scales, and also, rapid object movement makes accurate tracking more complicated, which may lead to potential errors and misidentifications, and a method and device for solving these problems can be provided.

[0023] In addition, according to the problem solving means of the present invention described above, two approaches are proposed for object detection, and an accuracy-first mode that can achieve a balance between high accuracy and speed required in a smartphone environment and a speed-first mode that can emphasize optimization of inference speed for tracking fast-moving objects can be provided.

[0024] FIG. 1 illustrates an exemplary environment in which an image recognition device according to some embodiments of the present disclosure may be applied.

[0025] FIG. 2 is a flowchart illustrating a method for recognizing and tracking an object according to an object recognition mode that an image recognition device can perform according to some embodiments of the present disclosure.

[0026] FIGS. 3 and 4 are exemplary diagrams of an architecture for a deep learning model corresponding to an accuracy priority mode according to some embodiments of the present disclosure.

[0027] FIG. 5 is an exemplary diagram of an architecture for a deep learning model corresponding to a speed priority mode according to some embodiments of the present disclosure.

[0028] FIG. 6 is an exemplary diagram of an architecture for a deep learning model for a process of assigning an ID in some embodiments of the present disclosure.

[0029] FIG. 7 is an exemplary diagram illustrating an example of an object recognized according to an accuracy priority mode according to some embodiments of the present disclosure.

[0030] FIG. 8 is an exemplary diagram illustrating an example of an object recognized in a speed priority mode according to some embodiments of the present disclosure.

[0031] FIG. 9 is a diagram of an exemplary computing device that may implement a device and / or system according to various embodiments of the present disclosure.

[0032] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the attached drawings. Advantages and features of the present disclosure, and methods for achieving them, will become clear with reference to the embodiments described in detail below together with the attached drawings. However, the technical idea of ​​the present disclosure is not limited to the following embodiments, but may be implemented in various different forms, and the following embodiments are provided only to complete the technical idea of ​​the present disclosure and to fully inform those skilled in the art of the present disclosure of the scope of the present disclosure, and the technical idea of ​​the present disclosure is defined only by the scope of the claims.

[0033] When assigning reference numerals to components in each drawing, it should be noted that identical components are assigned the same numerals whenever possible, even if they appear on different drawings. Furthermore, when describing the present disclosure, if a detailed description of a related known configuration or function is deemed likely to obscure the gist of the present disclosure, such detailed description will be omitted.

[0034] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in a meaning that can be commonly understood by a person of ordinary skill in the art to which this disclosure belongs. In addition, terms defined in commonly used dictionaries shall not be interpreted ideally or excessively unless explicitly specifically defined. The terminology used herein is for the purpose of describing embodiments and is not intended to limit the present disclosure. In this specification, the singular also includes the plural unless specifically stated otherwise in the phrase.

[0035] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the present disclosure. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When it is described that a component is "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.

[0036] The terms "comprises" and / or "comprising" as used in the specification do not exclude the presence or addition of one or more other components, steps, operations and / or elements.

[0037] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0038] In addition, when describing the components of the present invention, terms such as first, second, A, B, (a), (b), etc. may be used. These terms are only to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. Throughout the specification, when a part is said to "include" or "have" a component, this does not mean that other components are excluded, but rather that other components can be further included, unless specifically stated otherwise. In addition, terms such as "part" and "module" described in the specification mean a unit that processes at least one function or operation, and this can be implemented by hardware, software, or a combination of hardware and software.

[0039]

[0040] An object in image data can be detected and tracked through a system including an image collection device (100) and an image recognition device (200) as shown in Fig. 1, and in the process, an operation can be performed by utilizing a deep learning model according to an object recognition mode.

[0041] Below, the operations of the components illustrated in Fig. 1 related to the operation of recognizing and classifying an image provided through the above-described system will be described in more detail.

[0042] FIG. 1 illustrates an example in which an image collection device (100) and an image recognition device (200) are connected via a network, but this is only for convenience of understanding, and the number of devices that can be connected to the network may vary.

[0043] Meanwhile, Fig. 1 merely illustrates a preferred embodiment for achieving the purpose of the present disclosure, and some components may be added or deleted as needed. Below, the components illustrated in Fig. 1 will be described in more detail.

[0044] The image recognition device (200) can recognize an object from at least one image data input from the image collection device (100) and classify it according to object classification criteria.

[0045] The image collection device (100) may be a device in which image data is stored. The image collection device (100) may communicate via a network. The network may be implemented as any type of wired / wireless network, such as a Local Area Network (LAN), a Wide Area Network (WAN), a mobile radio communication network, Wibro (Wireless Broadband Internet), etc. The image collection device (100) may continuously acquire image data via the network, and such image data may be stored in the image collection device (100) and transmitted to the image recognition device (200). In addition, such an image collection device (100) may correspond to a camera sensor, a lidar, a lidar sensor, etc.

[0046] The image data described in the present invention is not limited to its storage format, and the image data may be image data including at least one object. Furthermore, the image data may also be image data that does not include an object. Furthermore, the image data may include not only still cut images but also video data that can be played back as a moving image.

[0047] The image recognition device (200) can recognize objects in an image and classify them by utilizing a deep learning algorithm based on various collected image data.

[0048] To avoid redundant description, the various operations performed by the image recognition device (200) will be described in more detail later with reference to the drawings below, FIG. 2.

[0049] Meanwhile, the image recognition device (200) may be implemented with one or more computing devices. For example, all functions of the image recognition device (200) may be implemented with a single computing device. As another example, the first function of the image recognition device (200) may be implemented with a first computing device, and the second function may be implemented with a second computing device. Here, the computing device may be, but is not limited to, a notebook, a desktop, a laptop, etc., and may include all types of devices equipped with computing functions. However, the image recognition device (200) may preferably be implemented with a high-performance, server-grade computing device. An example of a computing device will be described with reference to FIG. 9. Furthermore, it goes without saying that the image collection device (100) may also be implemented with one or more computing devices.

[0050] In some embodiments, components included in an environment to which the image recognition device (200) is applied may communicate via a network. The network may be implemented as any type of wired / wireless network, such as a local area network (LAN), a wide area network (WAN), a mobile radio communication network, or Wibro (Wireless Broadband Internet).

[0051] Meanwhile, the environment illustrated in FIG. 1 illustrates an image collection device (100) and an image recognition device (200) being connected via a network, but it should be noted that the scope of the present disclosure is not limited thereto, and the image collection devices (100) may also be connected P2P (Peer to Peer).

[0052] So far, with reference to FIG. 1, an exemplary environment in which an image recognition device (200) according to some embodiments of the present disclosure can be applied has been described. Hereinafter, with reference to the drawings including FIG. 2 and below, methods according to various embodiments of the present disclosure will be described in detail.

[0053] Each step of the methods described below may be performed by a computing device. In other words, each step of the methods may be implemented by one or more instructions executed by a processor of the computing device. All steps included in these methods may be performed by a single physical computing device, but first steps of the method may be performed by a first computing device, and second steps of the method may be performed by a second computing device.

[0054] In the following FIG. 2, the explanation will continue assuming that each step of the methods is performed by the image recognition device (200) illustrated in FIG. 1. However, for convenience of explanation, the description of the operating entity of each step included in the methods may be omitted.

[0055] Additionally, it should be noted that in the present invention, machine learning, deep learning, artificial intelligence, and AI algorithm are all terms that can be used identically, and can generally refer to all machine learning algorithms within the scope that can be understood by a person skilled in the art.

[0056]

[0057] FIG. 2 is a flowchart illustrating a method for recognizing and tracking an object according to an object recognition mode that an image recognition device can perform according to some embodiments of the present disclosure.

[0058] In step S100, the image recognition device (200) can load image data. The image recognition device (200) can acquire image data. The image recognition device (200) can receive image data through the image collection device (100), and the image recognition device (200) can utilize the image data to train a deep learning model.

[0059] The above image data may be a 2D image or a 3D image containing an object within the image, and its storage format is not limited to one example. Accordingly, the image data may be image data stored in various storage formats such as .jpg, .png, .jpeg, and .mov.

[0060] Additionally, image data may include at least one object within the image data, or may be image data that does not include any object. Therefore, it should be noted that the image data in the present invention may be image data related to drawings, photographs, etc. that can be obtained by a person skilled in the art, and is not limited to a specific image.

[0061] In step S200, the image recognition device (200) may preprocess image data. The preprocessing process may include performing a labeling process on image data acquired to train a deep learning model.

[0062] The image recognition device (200) can output object recognition values ​​within image data through a trained deep learning model. Alternatively, according to an example, the image data may be data generated by performing a labeling process at the time the image data is loaded from the image collection device (100) to the image recognition device (200). Accordingly, when data that does not require a preprocessing process is loaded, the above step S200 may be omitted.

[0063] Here, the labeling process refers to the process of generating data for training a deep learning model, and refers to the process of generating data in which bounding boxes of areas where objects are located within image data are marked and information about the objects within the bounding boxes is labeled. For example, if the image data includes a dog or a cat, the labeling data may refer to the process of generating data in which at least one bounding box of an area where a dog or cat is located within the image is marked and a dog or cat is described within each of the at least one bounding box.

[0064] Through this process, the image recognition device (200) can perform a process of learning a deep learning model through a preprocessing process before inputting image data into the deep learning model.

[0065] In step S300, the image recognition device (200) can learn at least one deep learning model.

[0066] The training process of a deep learning model can be performed using image data that has undergone a preprocessing process. The preprocessed image data may be an image containing an object. The preprocessed image data may be image data in which a bounding box is drawn around an area where an object is expected to be present, and labeling information, which is information about what type of object the object is, is included. For example, the preprocessed image data may be data in which a bounding box is drawn around the area occupied by the dog in the image data, and labeling information such as “puppy,” “dog,” or “dog” is included in the bounding box, if there is a dog in the image data. The deep learning model can input preprocessed image data, output an output value, and learn by comparing the preprocessed image data with the output value. According to one example, if the preprocessed image data includes labeling data, the output value may be a value including a probability value regarding the labeling data.

[0067] Deep learning models can be trained using backpropagation. Backpropagation is a learning method that reflects the error between preprocessed image data and output values ​​back into the nodes of the deep learning model. Deep learning models can learn by minimizing the error between preprocessed image data and output values.

[0068] The image recognition device (200) can input image data into the deep learning model through the trained deep learning model and output an object recognition value. The image recognition device (200) can input image data into the trained deep learning model and output a recognition value regarding an object included in the image data, and the recognition value may be a value expressed as a probability value. In this case, the recognition value may be a value that expresses a value for the probability corresponding to the object by classifying what kind of object the object in the image data is. Therefore, the object recognition value may mean an output value output by the deep learning model in relation to image recognition in the deep learning model.

[0069] In step S400, the image recognition device (200) can set an object recognition mode. In the present invention, the image recognition device (200) can set at least one object recognition mode and recognize an object by utilizing a corresponding deep learning model according to the mode. The object recognition mode described above in the present invention may be a speed priority mode and an accuracy priority mode, and the speed priority mode may mean a mode that gives priority to the speed of object recognition and calculates, and the accuracy priority mode may mean a mode that gives priority to the accuracy of object recognition and calculates. In step S500, the image recognition device (200) can recognize an object by utilizing a deep learning model corresponding to the object recognition mode according to the object recognition mode.

[0070] Below, the architecture of a deep learning model for accuracy priority mode is described in detail through Figures 3 and 4.

[0071]

[0072] FIGS. 3 and 4 are exemplary diagrams of an architecture for a deep learning model corresponding to an accuracy priority mode according to some embodiments of the present disclosure.

[0073] Looking at Figures 3 and 4, the accuracy-first mode can be designed to achieve faster performance than the YOLOv8s model while minimizing potential accuracy degradation. It can be derived from the Path Aggregation Network (PANet), which serves as an extension of the Feature Pyramid Network (FPN) architecture deployed within the YOLOv8 context. This extension involves an additional expansion step on the reduced results of the FPN, which can improve inference accuracy compared to utilizing the standard FPN. Furthermore, this mode integrates a block rooted in the CSPNet application into ResNet, which is constructed through a modified transformation of the CSPNet architecture. While CSPNet divides the input data into two segments, connecting one segment to the original network and concatenating the output with the other, a designated block splits the input into two segments. One segment is connected to the original network with downsampling convolutional layers, and the other segment is interfaced with a max-pooling layer. The results of the two sections are then combined, and this structure can be combined with Residual Blocks inspired by ResNet to form CSPResidualBlock.

[0074] In the Neck segment, which acts as an intermediary between the Backbone and the Head, the mode starts with SPPCSP (Spatial Pyramid Pool, SPP with CSPNet) and switches to using ResidualBlocks instead of CSPResidualBlocks during scaling. This modification is due to the limitation of CSPResidualBlocks within this mode to effectively accommodate the difference between the input and output sizes. As the mode progresses through the upscaling phase, the number of filters may decrease. The CSPResNet architecture doubles the number of filters compared to the input, which poses a challenge when trying to match or halve the output size while maintaining a consistent input. This adjustment reduces the overall number of filters across the architecture, which may result in loss of important features. Consequently, after the upscaling phase is completed, subsequent downscaling incorporates CSPResidualBlocks. The increased number of filters at this stage can facilitate the effective application of this architectural structure.

[0075] Below, we will describe the architecture of a deep learning model corresponding to the speed priority mode using Fig. 5.

[0076]

[0077] FIG. 5 is an exemplary diagram of an architecture for a deep learning model corresponding to a speed priority mode according to some embodiments of the present disclosure.

[0078] As shown in Figure 5, models designed for speed-first mode are characterized by extremely fast inference speeds without significantly compromising accuracy. This particular mode simplifies the architecture by reducing the number of extraction layers and filters compared to accuracy mode.

[0079] Unlike the accuracy-first mode shown in Figure 5, the current mode is characterized by having only two outputs, which simplifies the output shape and maintains a consistent output size standard. The accuracy-first mode produces an output shape of (1, 7, 8400) based on a 640x640 pixel input size. Conversely, the speed-first mode significantly reduces the number of coordinates in the output array by more than four times, resulting in a shape of (1, 7, 2000) based on the same input size. This reduction not only simplifies processing and improves recognition speed, but also reduces computational overhead. Looking at Figure 5, a degree of simplification can be seen compared to the accuracy-first mode, where a single CSPResidualBlock is replaced with a downsampling convolutional layer. The head section is designed to be simpler to achieve the desired two outputs, and the ResidualBlock remains unused during expansion but is used once during the reduction process. This configuration can facilitate achieving dual outputs.

[0080] Returning to FIG. 2 again, in step S600, the image recognition device (200) can track the recognized object and assign an object ID corresponding to the object.

[0081] Because there's no built-in mechanism for assigning identification (ID) tags to individual object features, it's essential to implement a custom system within TensorFlow that can efficiently assign IDs. This allows for effective tracking and identification of specific objects. These unique IDs can be used to distinguish between various objects, such as "Person 0" and "Person 1."

[0082] The object ID assignment process consists of two separate steps. First, feature extraction is performed using a Convolutional Neural Network (CNN). Then, features are extracted using a CNN model called PlainNet, which is trained on a dataset containing humans, dogs, and cats. This CNN model is illustrated in Figure 6.

[0083]

[0084] FIG. 6 is an exemplary diagram of an architecture for a deep learning model for a process of assigning an ID in some embodiments of the present disclosure.

[0085] Figure 6 is a diagram visually illustrating the structure of a deep learning model related to the ID assignment process. The image recognition device (200) may perform feature extraction to assign an object ID, and then evaluate the similarity of feature maps to reassign an existing ID or assign a new ID. More specifically, if a feature map shows a similarity of 50% or more with an existing feature map, the same ID associated with the feature map should be reassigned. Conversely, if the similarity is lower than a specified threshold, a new ID is assigned, and the feature map can be stored in association with the newly assigned ID.

[0086] Returning to Figure 2 again, in step S600, the image recognition device (200) can track an object to which an ID is assigned.

[0087] At this time, the API of the cradle installed in the image recognition device (200) can be used to connect to the cradle device and track the physical object. The API code for controlling the cradle in the image recognition device (200) can be performed based on the code provided by Pivo, and the cradle control functions can include 'turnLeft', 'turnRight', 'turnLeftContinuous', and 'turnRightContinuous'.

[0088] To implement smooth, continuous motion, you can use the 'turnLeftContinuous' and 'turnRightContinuous' functions. Both will continue to rotate at the speed set before the function call if no speed value is provided as an argument. The rotation speed value does not represent angular velocity, but rather the time it takes for the rotation to complete a full 360-degree cycle. For example, setting the rotation speed to 10 would mean that the object completes a full 360-degree rotation in 10 seconds, and this relationship can be expressed as Equation 1.

[0089] [Mathematical Formula 1]

[0090]

[0091] In Equation 1, the variable "rt" (Rotational Time) represents the time required for a complete 360-degree rotation and is used as an input for the cradle control function. It serves as an indicator of the time it takes for the cradle to complete a full rotation. Similarly, the variable v_actual represents the cradle's real-time rotational speed, measured in degrees per second. Therefore, if the cradle's actual rotational speed decreases, "rt" increases. Considering its characteristic of decreasing with acceleration, the rotational time can be easily calculated.

[0092] The rotation speed of the cradle is defined as the duration required to complete a full 360° rotation and is designed to provide adjustable cradle rotation speed, allowing for rapid and gradual cradle rotations depending on the proximity of an object.

[0093] The degree of object displacement relative to the camera's center of focus can be determined based on the object's x-axis coordinate relative to the camera center. Furthermore, the cradle's rotation time is intentionally configured to decrease as the object approaches the center of the camera image. This feature is crucial to prevent situations where the cradle cannot stop immediately even if the object is stationary, causing the object to move laterally due to being out of alignment with the center of the camera image. Consequently, by leveraging this unique characteristic of the cradle, the following formula can be used to calculate the rotation time based on the distance between the object and the center of the screen.

[0094] [Equation 2]

[0095]

[0096]

[0097] [Equation 3]

[0098]

[0099] Mathematical equation (2) describes the preprocessing steps involved in determining the x-axis center coordinate of an object. In this equation, 'OP' (Object Position) represents the x-axis coordinate of the object. It is important to emphasize that in this experiment, the object position is normalized within the range of 0 to 1, measured in pixels, for computational efficiency, and the distance from the center of the screen to the object is expressed as 'MD' (Moving Distance). Specifically, 'P=0.5' is designated when the object is aligned to the center of the camera, 'P=1' is expressed when the object is located at the rightmost edge of the camera, and 'P=0' indicates that the object is at the leftmost edge of the camera. Since the cradle used in the present invention does not involve tilting, the vertical (y-axis) movement of the object was not considered.

[0100] In Equation 2, MD is designed to return 0 when the object's position (OP) is in the range of 0.4 to 0.6, which is near the center of the screen (OP=0.5). Here, MD represents the distance from the center of the screen to the object and is a standardized value as explained above. Since the distance from the center of the screen is 0.5, which is the minimum, it indicates that there is no need to move the cradle even if the object is outside this area. For positions outside this range, 0.5 is subtracted from OP and the result is multiplied by 2 to ensure that it falls within the range of -1 and 1. When MD approaches 1, this indicates that the object is at the far right edge. In this case, to reposition the object toward the center of the camera, the cradle must be rotated more quickly to the right. Conversely, when MD approaches -1, this indicates that the object is at the far left edge. In this case, to reposition the object at the far left edge toward the center of the camera, the cradle must be moved in the opposite direction.

[0101] In Equation 3, RT (Rotational Time) is recalculated as the time required for the cradle to complete a 360° rotation based on the MD value. The cradle control part uses the API provided by the PIVO manufacturer, and the RT value ranges from 6 to infinity. When the RT value is 6, the cradle rotates at maximum speed, and as RT increases toward infinity, the cradle's rotation speed decreases toward a value of 0. The design ensures that the RT value approaches 6 as MD approaches 1 or -1. Conversely, as the object gets closer to the center of the screen, MD converges to 0 and RT approaches infinity, resulting in a slower rotation speed. In addition, to prevent a ZeroDivisionError when the preprocessed MD value is 0, RT is set to return 0 in such cases.

[0102] Next, the acquired RT value can be applied to rotate the cradle. If RT is positive, the cradle rotates to the right, and if it is negative, it rotates to the left. If RT is 0, the cradle remains fixed and no rotation occurs.

[0103] Below, an example screen in which object recognition is implemented using Figures 7 and 8 is described.

[0104]

[0105] FIGS. 7 and 8 are exemplary drawings illustrating an example of an object recognized according to an accuracy priority mode and an example of an object recognized according to a speed priority mode, according to some embodiments of the present disclosure.

[0106] Figures 7 and 8 provide an overview of the object recognition components within a visual display application, and the object recognition screen can provide several key features. First, the two models of the application use different input sizes: the accuracy-first mode operates with an input size of 640x640, while the speed-first mode can use an input size of 320x320. Second, users and content creators can choose between the accuracy-first and speed-first modes. Third, this screen integrates an object selection function, allowing users to specify the objects they wish to recognize. Users can choose to detect people, dogs, cats, or all objects.

[0107] Additionally, the screen displays the assigned ID for recently recognized objects. This ID is determined by analyzing the identified object's characteristics, such as size, shape, and color. The assignment process ensures that the same object consistently receives the same ID, facilitating object tracking over time. Object selection and ID assignment in applications are crucial for accurate and reliable object recognition, and by allowing users to select the objects they wish to recognize, the application can be customized to meet their specific needs.

[0108] Below, a system to which an embodiment of the present invention can be applied will be described using FIG. 9.

[0109]

[0110] FIG. 9 is a diagram of an exemplary computing device that may implement a device and / or system according to various embodiments of the present disclosure.

[0111] A computing device (1500) may include one or more processors (1510), a bus (1550), a communication interface (1570), a memory (1530) for loading a computer program (1591) to be executed by the processor (1510), and a storage (1590) for storing the computer program (1591). However, only components related to the embodiment of the present disclosure are illustrated in FIG. 9. Therefore, a person skilled in the art to which the present disclosure pertains may recognize that other general components may be included in addition to the components illustrated in FIG. 9.

[0112] The processor (1510) controls the overall operation of each component of the computing device (1500). The processor (1510) may include a central processing unit (CPU), a microprocessor unit (MPU), a microcontroller unit (MCU), a graphics processing unit (GPU), or any other type of processor well known in the art of the present disclosure. In addition, the processor (1510) may perform operations for at least one application or program for executing a method according to embodiments of the present disclosure. The computing device (1500) may include one or more processors.

[0113] The memory (1530) stores various data, commands, and / or information. The memory (1530) may load one or more programs (1591) from the storage (1590) to execute a method according to embodiments of the present disclosure. The memory (1530) may be implemented as a volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0114] The bus (1550) provides communication between components of the computing device (1500). The bus (1550) may be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0115] The communication interface (1570) supports wired and wireless Internet communication of the computing device (1500). Furthermore, the communication interface (1570) may support various communication methods other than Internet communication. To this end, the communication interface (1570) may be configured to include a communication module well known in the technical field of the present disclosure.

[0116] According to some embodiments, the communication interface (1570) may be omitted.

[0117] Storage (1590) can non-temporarily store one or more programs (1591) and various data.

[0118] Storage (1590) may be configured to include non-volatile memory such as Read Only Memory (ROM), Erasable Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which the present disclosure pertains.

[0119] The computer program (1591) may include one or more instructions that, when loaded into the memory (1530), cause the processor (1510) to perform methods / operations according to various embodiments of the present disclosure. That is, the processor (1510) may perform the methods / operations according to various embodiments of the present disclosure by executing the one or more instructions.

[0120] Various embodiments of the present disclosure and effects according to the embodiments have been described with reference to FIGS. 1 through 9 so far. The effects according to the technical concept of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description in the specification.

[0121] The technical idea of ​​the present disclosure described with reference to FIGS. 1 to 9 so far can be implemented as a computer-readable code on a computer-readable medium. The computer-readable recording medium can be, for example, a removable recording medium (CD, DVD, Blu-ray disc, USB storage device, removable hard disk) or a fixed recording medium (ROM, RAM, computer-attached hard disk). The computer program recorded on the computer-readable recording medium can be transmitted to another computing device via a network such as the Internet and installed on the other computing device, thereby allowing it to be used on the other computing device.

[0122] Although all components constituting the embodiments of the present disclosure have been described as being combined or operated in combination as one, the technical concept of the present disclosure is not necessarily limited to such embodiments. That is, within the scope of the present disclosure, all components may be selectively combined and operated one or more times.

[0123] Although operations are depicted in the drawings in a particular order, this should not be understood to imply that the operations must be performed in the particular order depicted, or in any sequential order, or that all depicted operations must be performed to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various components in the embodiments described above should not be understood to imply that such separation is absolutely necessary, and it should be understood that the program components and systems described may generally be integrated together into a single software product or packaged into multiple software products.

[0124] Although the embodiments of the present disclosure have been described with reference to the attached drawings, those skilled in the art will appreciate that the present disclosure can be implemented in other specific forms without changing the technical idea or essential features thereof. Therefore, it should be understood that the embodiments described above are exemplary in all respects and not restrictive. The scope of protection of the present disclosure should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the technical ideas defined by the present disclosure.

Claims

1. A method performed by an image recognition device, Step of loading image data; Step of preprocessing image data; A step of training at least one deep learning model by utilizing the above preprocessed image data; Step for setting the mode for recognizing objects; A step of recognizing an object by utilizing a deep learning model corresponding to the object recognition mode according to the object recognition mode; and A step of tracking a recognized object and assigning an object ID; comprising; Real-time moving object tracking method using cradle head servo motor.

2. In paragraph 1, The mode for recognizing the above object is characterized by including an accuracy priority mode and a speed priority mode. Real-time moving object tracking method using cradle head servo motor.

3. In paragraph 2, The deep learning model corresponding to the above accuracy priority mode is built using the CSPNet structure integrated with ResNet based on the YOLOv8 architecture. Real-time moving object tracking method using cradle head servo motor.

4. In paragraph 2, The deep learning model corresponding to the above speed-first mode is YOLOv5n, which is built using the shallowest architecture. Real-time moving object tracking method using cradle head servo motor.

5. In paragraph 1, The step of tracking the recognized object and assigning an object ID is characterized in that if the feature map similarity between the object features in the current frame and the object features in the previous frame exceeds 50% based on the feature map, the same object ID is assigned. Real-time moving object tracking method using cradle head servo motor.

6. In paragraph 1, The step of recognizing an object by using a deep learning model corresponding to the object recognition mode is as follows: It is characterized in that the x-axis coordinate position of the object is calculated by utilizing the mathematical expression 1 in the cradle included in the image recognition device. [Mathematical formula 1] The above MD is the distance from the center of the screen to the object, and OP represents the x-axis coordinate of the object. Real-time moving object tracking method using cradle head servo motor.

7. In paragraph 6, The step of recognizing an object by utilizing a deep learning model corresponding to the object recognition mode according to the object recognition mode is characterized by calculating the time required for the cradle to rotate 360 ​​degrees based on the MD value by utilizing mathematical expression 2. [Mathematical formula 2] The above RT is the time required for the cradle to rotate 360 ​​degrees. Real-time moving object tracking method using cradle head servo motor.

8. Processor; network interface; memory; and A computer program loaded into the above memory and executed by the above processor, The above processor, Instructions for loading image data; Instructions for preprocessing image data; Instructions for training at least one deep learning model using the above preprocessed image data; An instruction that sets the mode for recognizing objects; Instructions for recognizing an object by utilizing a deep learning model corresponding to the object recognition mode according to the object recognition mode; and Instructions for tracking recognized objects and assigning object IDs; including; Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

9. In paragraph 8, The mode for recognizing the above object is characterized by including an accuracy priority mode and a speed priority mode. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

10. In paragraph 9, The deep learning model corresponding to the above accuracy priority mode is built using the CSPNet structure integrated with ResNet based on the YOLOv8 architecture. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

11. In paragraph 9, The deep learning model corresponding to the above speed-first mode is YOLOv5n, which is built using the shallowest architecture. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

12. In paragraph 8, The instruction for tracking the recognized object and assigning an object ID is characterized in that if the feature map similarity between the object features in the current frame and the object features in the previous frame exceeds 50% based on the feature map, the same object ID is assigned. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

13. In paragraph 8, Depending on the object recognition mode, an instruction is provided to recognize an object using a deep learning model corresponding to the object recognition mode. It is characterized in that the x-axis coordinate position of the object is calculated by utilizing the mathematical expression 1 in the cradle included in the image recognition device. [Mathematical formula 1] The above MD is the distance from the center of the screen to the object, and OP represents the x-axis coordinate of the object. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

14. In paragraph 13, Depending on the object recognition mode, the instruction for recognizing an object using a deep learning model corresponding to the object recognition mode is characterized by calculating the time required for the cradle to rotate 360 ​​degrees based on the MD value using mathematical expression 2. [Mathematical formula 2] The above RT is the time required for the cradle to rotate 360 ​​degrees. Image recognition device capable of real-time moving object tracking using a cradle head servo motor.

Citation Information

Patent Citations

  • Method and device for providing club path image

    KR1020250022295A

  • Solar Cell Manufacturing Appratus and Solar Cell Manufacturing Method

    KR1020250056007A

  • Vision-based safety monitoring and / or activity analysis

    US20230072434A1

  • System and method for surveillance of goods

    US20230245460A1

  • Method and system for automated evaluation of animals

    US20230342902A1