Method and apparatus for detecting object for autonomous navigation
The object detection method and device for autonomous navigation address navigation challenges by using modified image processing and learning techniques to enhance object detection accuracy and safety in recreational vessels.
Patent Information
- Application Number
- PCT/KR2025/011950
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-07
- Filing Date
- 2025-08-07
- Publication Date
- 2026-02-12
AI Technical Summary
Recreational vessel operators face difficulties in navigating and berthing due to inexperience and environmental disturbances, hindering general public access to operating such vessels.
An object detection method and device for autonomous navigation using a camera to collect images, modify them, and input into an object detection model for accurate detection of objects around the ship, with a learning process that generates high-quality models through abundant data and adaptive learning.
Enables accurate object detection for safe ship operation by maintaining model performance even with new data inputs, enhancing navigation capabilities.
Smart Images

Figure KR2025011950_12022026_PF_FP_ABST
Abstract
Description
Object detection method and device for autonomous navigation
[0001] The present invention relates to an object detection method and device for autonomous navigation.
[0002] Typically, users of small vessels operate the steering wheel and throttle to navigate, berth, or unberth their vessels. However, due to the inexperience of users of recreational vessels and the influence of environmental disturbances such as currents and winds, vessels often experience difficulties navigating, especially when berthing or unberthing in confined spaces. These difficulties hinder the general public's access to operating or manoeuvring recreational vessels.
[0003] Accordingly, various technologies to assist ship operation are being actively developed.
[0004] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired in the process of deriving the present invention, and cannot necessarily be considered as publicly known technology disclosed to the general public prior to the application for the present invention.
[0005] The purpose of the present disclosure is to provide an object detection method and device for autonomous navigation. The problems addressed by the present disclosure are not limited to those mentioned above. Other problems and advantages of the present disclosure, not mentioned above, can be understood through the following description and will be more clearly understood through the embodiments of the present disclosure. Furthermore, it will be appreciated that the problems and advantages addressed by the present disclosure can be realized by the means and combinations thereof set forth in the claims.
[0006] A first aspect of the present disclosure may provide an object detection method for autonomous navigation, including the steps of: receiving an original image collected by a camera; generating a modified image by modifying the original image; inputting the modified image into an object detection model for learning the object detection model; and detecting an object existing around a ship based on the object detection model.
[0007] A second aspect of the present disclosure provides an object detection device for autonomous navigation, comprising: a memory having at least one program stored therein; and a processor that operates by executing the at least one program; wherein the processor receives an original image collected by a camera, modifies the original image to generate a modified image, inputs the modified image into an object detection model for learning an object detection model, and detects an object existing around a ship based on the object detection model.
[0008] A third aspect of the present disclosure can provide a computer-readable recording medium having recorded thereon a program for executing the method according to the first aspect on a computer.
[0009] According to various embodiments of the present disclosure, abundant learning data can be formed, and as a result, a high-quality object detection model can be generated.
[0010] In particular, since images containing objects that are difficult to capture by cameras, such as drowning people or glaciers, can be generated, multi-faceted learning of object detection models can be achieved.
[0011] Additionally, since newly generated learning data is generated based on previously used learning data, the object detection model can perform learning only on the changed parts, so even if new data is input, the performance of the object detection model may not deteriorate during the learning process.
[0012] As a result, objects can be accurately detected based on a model with excellent performance, enabling safe ship operation.
[0013] Figure 1 is a diagram showing an example of an operation system including a ship and a control server.
[0014] FIG. 2 is a flowchart of an object detection method according to one embodiment of the present disclosure.
[0015] Figure 3 is a conceptual diagram illustrating an inference step of an artificial neural network according to one embodiment.
[0016] Figure 4 is a conceptual diagram showing a learning step of an artificial neural network according to one embodiment.
[0017] FIG. 5 is a conceptual diagram for explaining the operation of a learning image generation unit according to one embodiment of the present disclosure.
[0018] FIG. 6 is a conceptual diagram for explaining the concept of setting data according to one embodiment of the present disclosure.
[0019] FIGS. 7A to 7E are drawings illustrating examples of deformed images according to embodiments of the present disclosure.
[0020] FIG. 8 is a flowchart illustrating a method for generating a deformed image according to another embodiment of the present disclosure.
[0021] FIG. 9 is a block diagram of a device according to one embodiment of the present disclosure.
[0022] According to one embodiment of the present invention for solving the above technical problem, a method for detecting an object for autonomous navigation comprises the steps of: receiving an original image collected by a camera; generating a modified image by modifying the original image; inputting the modified image into an object detection model for learning the object detection model; and detecting an object existing around a ship based on the object detection model.
[0023] In the above method, the step of generating the modified image includes a step of detecting one or more objects included in the original image; and the modified image can be generated based on the one or more objects.
[0024] In the above method, the transformed image can be generated based on at least one of a process of moving at least one of the one or more objects, a process of replacing at least one of the one or more objects, or a process of adding an object other than the one or more objects.
[0025] In the above method, the method may further include a step of deleting at least some of one or more objects included in the original image to create a base image and storing the base image in a database.
[0026] In the above method, the step of generating the transformed image may include the step of loading a base image from the database; and the step of adding a third object to the base image.
[0027] In the above method, the step of generating the transformed image may include a step of generating labeling data for each of one or more objects included in the transformed image.
[0028] In the above method, the method further includes a step of receiving deformation type setting data and object type setting data input by a user; and the generation of the deformation image can be based on the deformation type setting data and the object type setting data.
[0029] In the above method, the method may further include a step of generating an information providing interface based on the detected object.
[0030] According to another embodiment of the present invention for solving the above technical problem, a device is an object detection device for autonomous navigation, comprising: a ship device, a memory having at least one program stored therein; and a processor operating by executing the at least one program; wherein the processor receives an original image collected by a camera, modifies the original image to generate a modified image, inputs the modified image into an object detection model for learning an object detection model, and detects an object existing around the ship based on the object detection model.
[0031] One embodiment of the present invention can provide a computer-readable recording medium having recorded thereon a program for executing the above method on a computer.
[0032] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments presented below, but can be implemented in various different forms, and it should be understood that it includes all transformations, equivalents, and substitutes included in the spirit and technical scope of the present invention. The embodiments presented below are provided to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the invention of the scope of the invention. In describing the present invention, if a detailed description of a related known technology is judged to obscure the gist of the present invention, the detailed description thereof will be omitted.
[0033] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0034] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a given function. Furthermore, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented by algorithms that execute on one or more processors. Furthermore, the present disclosure may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.
[0035] Additionally, the connecting lines or connecting members between components depicted in the drawings are merely exemplary representations of functional connections and / or physical or circuit connections. In an actual device, connections between components may be represented by various functional connections, physical connections, or circuit connections that may be replaced or added.
[0036] Additionally, terms including ordinal numbers, such as "first" or "second," used in the specification may be used to describe various components, but the components should not be limited by the terms. The terms may be used to distinguish one component from another.
[0037] Figure 1 is a conceptual diagram for explaining an autonomous navigation system according to the present disclosure.
[0038] Referring to FIG. 1, the autonomous navigation system (100) may include a ship's control device (110), an EIU (Engine Interface Unit 120), an Autonomous Navigation Processor (130, hereinafter, an autonomous navigation processing unit), and an engine (140).
[0039] The steering device (110) may include at least a portion of a throttle lever, a steering wheel, and a joystick. However, the present invention is not limited thereto, and the steering device (110) may include other ship devices. The steering device (110) may be referred to as a helm station according to an embodiment.
[0040] The EIU (120) may refer to a device that acquires a signal (S1) from a steering device (110) included in a ship through a communication network within the ship and transmits it to the engine (140). The signal (S1) may include a message or protocol of several steering devices (110) included in the ship. The EIU (120) may transmit the signal (S1) of the steering device (110) as is to the engine (140), or may convert the signal (S1) of the steering device (110) and then transmit the converted signal (S2) to the engine (140). The present invention is not limited thereto.
[0041] The EIU (120) may be a device that enables switching between the autonomous navigation mode (or automatic navigation mode) and the manual navigation mode of the ship by transmitting or injecting a signal (S1) of the steering device (110). The EIU (120) may determine whether the current navigation mode of the ship is the autonomous navigation mode or the manual navigation mode based on a control command received from the autonomous navigation processing device (130). The EIU (120) may also receive the navigation status of the ship determined from the autonomous navigation processing device (130). In this case, the EIU (120) may operate based on the determined navigation status. According to an embodiment, the EIU (120) may determine the autonomous navigation mode by detecting a control value output from the steering device (110) even when the user does not control the steering devices (110).
[0042] The EIU (120) can be connected to the ship's steering device (110) and engine (140) via an internal communication network. The internal communication network can include a CAN (Controller Area Network). Here, CAN can refer to an internal communication network of a ship, automobile, etc. that can perform data transmission between ECUs (Engine Control Units), control of various steering devices (110), control of a system, etc., and is not limited thereto. The internal communication network can refer to any communication network that can transmit data between the ship's steering device (110) and the engine (140). According to an embodiment, a user can check the status of the linkage between the steering devices (110), the autonomous navigation processing unit (130), and the engine (140) by using the EIU (120).
[0043] The autonomous navigation processing unit (130) may be a device that processes control commands for controlling the ship's steering device (110) when the ship is in autonomous navigation mode. In other words, even when the user does not operate the steering device (110), the autonomous navigation processing unit (130) may generate commands for controlling the ship's steering device (110). In addition, the autonomous navigation processing unit (130) may transmit the ship's navigation status to the EIU (120) so that the EIU (120) may operate based on the ship's navigation status.
[0044] The engine (140) may be a device that operates based on control commands from the ship's steering devices (110). For example, when the ship is in manual navigation mode, the engine (140) may operate based on user input operating the steering device (110). As another example, when the ship is in autonomous navigation mode, the engine (140) may operate based on control commands received from the autonomous navigation processing device (130).
[0045] FIG. 2 is a flowchart of an object detection method according to one embodiment of the present disclosure.
[0046] Each step of the object detection method illustrated in FIG. 2 may be performed by an object detection device or an autonomous navigation processing device (130) including an object detection device. Specifically, each step of the object detection method illustrated in FIG. 2 may be performed by a processor of the object detection device or a processor of an autonomous navigation processing device (130) including an object detection device.
[0047] Referring to FIG. 2, at step 210, the processor may receive an original image collected by the camera.
[0048] Referring to FIG. 2, at step 220, the processor may transform the original image to generate a transformed image.
[0049] In one embodiment, step 220 may include detecting one or more objects contained in the original image.
[0050] In one embodiment, the transformed image may be generated based on one or more objects.
[0051] In one embodiment, the transformed image may be generated based on at least one of: moving at least one of the one or more objects, replacing at least one of the one or more objects, or adding one or more objects to another object.
[0052] In one embodiment, the processor may generate a base image by deleting at least some of one or more objects included in the original image, and store the base image in a database.
[0053] In one embodiment, step 220 may include loading a base image from a database and adding a third object to the base image.
[0054] In one embodiment, step 220 may include generating labeling data for each of one or more objects included in the transformed image.
[0055] In one embodiment, the processor may receive transformation type setting data and object type setting data input by a user.
[0056] In one embodiment, step 220 may be based on transformation type setting data and object type setting data.
[0057] Referring to FIG. 2, at step 230, the processor may input the transformed image into the object detection model for learning the object detection model.
[0058] In one embodiment, the processor can detect objects present in the vicinity of the ship based on an object detection model.
[0059] In one embodiment, the processor may generate an information providing interface based on the detected object.
[0060] Hereinafter, a method for training an object detection model and an object detection method according to various embodiments performed by an object detection device of the present disclosure are described in detail.
[0061] Figure 3 is a conceptual diagram illustrating an inference step of an artificial neural network according to one embodiment.
[0062] Input data (301) may be input to an artificial neural network (300), and output data (302) may be output by the artificial neural network (300). The types of each of the input data (301) and the output data (302) may vary based on the purpose for which the artificial neural network (300) is configured. The input data (301) and the output data (302) of the artificial neural network (300) may vary depending on which data the artificial neural network (300) is trained to output. In addition, the accuracy of the output data (302) may vary depending on the degree of learning of the artificial neural network (300).
[0063] In the present disclosure, an artificial neural network (300) may be included in an autonomous navigation processing device (130). Specifically, the autonomous navigation processing device (130) may include an object detection model that detects an object by processing an image collected by a camera mounted on a ship, and the object detection model may include an artificial neural network (300). In this case, input data (301) may be an image collected by the camera, and output data (302) may include information about the detected object. For example, the output data (302) may include a bounding box indicating the location of the detected object, a class of the detected object, a probability that the detected object actually corresponds to the class, etc.
[0064] Figure 4 is a conceptual diagram showing a learning step of an artificial neural network according to one embodiment.
[0065] The artificial neural network (400) illustrated in FIG. 4 can be understood to correspond to the artificial neural network (300) illustrated in FIG. 3.
[0066] An artificial neural network (400) can receive learning data (401) as input. The learning data (401) may be input data provided so that the artificial neural network (400) can learn a pattern. The artificial neural network (400) can process the learning data (401) to generate output data (402). The output data (402) may be a result predicted by the artificial neural network (400).
[0067] The output data (402) of the artificial neural network (400) can be compared with labeling data (403). The labeling data (403) can mean the correct answer corresponding to the learning data (401). Based on the comparison of the output data (402) and the labeling data (403), an error, i.e., the difference between the predicted result and the correct answer, can be calculated.
[0068] The error can be minimized based on the backpropagation algorithm. Specifically, the weights and biases of the artificial neural network (400) can be adjusted using the backpropagation algorithm based on the error.
[0069] In the present disclosure, the training data (401) may be images. As will be described below, the training data (401), which is an image, may include not only images collected by a camera, but also modified images generated according to various embodiments of the present disclosure. In the present disclosure, the output data (402), similar to the output data (302) described above, may include information about a detected object. The labeling data (403) may include accurate information about the objects included in the training data (401).
[0070] FIG. 5 is a conceptual diagram for explaining the operation of a learning image generation unit according to one embodiment of the present disclosure.
[0071] The autonomous navigation processing device (130) of the present disclosure may include a learning image generation unit (500). In another embodiment, a server may include the learning image generation unit (500). The server may be installed onboard or offboard. The autonomous navigation processing device (130) and / or the server may update the object detection model (510) using the transformed image (502) of the learning image generation unit (500).
[0072] In one embodiment, the learning image generation unit (500) may receive an original image (501) and generate a transformed image (502). The learning image generation unit (500) may be understood as a functional block of an autonomous navigation processing device (130) that performs an operation of receiving an original image (501) and generating a transformed image (502).
[0073] In one embodiment, the transformed image (502) may be an image generated by transforming the original image (501). In one embodiment, the learning image generation unit (500) may generate the transformed image (502) by transforming the original image (501).
[0074] The operations performed by the learning image generation unit (500) of the present disclosure will be described in detail later.
[0075] The autonomous navigation processing device (130) of the present disclosure may include an object detection model (510). The object detection model (510) may correspond to the artificial neural network (300) of FIG. 3 or the artificial neural network (400) of FIG. 4 described above.
[0076] In one embodiment, the object detection model (510) may be trained based on a transformed image (502) and labeling data (503). The training step of the artificial neural network may be understood with reference to FIG. 4. The transformed image (502) may be included in the training data (401) of FIG. 4, and the labeling data (503) may correspond to the labeling data (403) of FIG.
[0077] In one embodiment, labeling data (503) may be generated by a learning image generation unit (500). The learning image generation unit (500) may generate labeling data (503) together with a transformed image (502). In this case, the labeling data (503) may relate to each of one or more objects included in the transformed image (502).
[0078] Hereinafter, various embodiments in which the learning image generation unit (500) generates a transformed image (502) are described. Through the transformed image (502) generated by the learning image generation unit (500), the object detection model (510) can be trained based on abundant learning data. The embodiments described below in terms of the operation of the learning image generation unit (500) can be understood as being performed by the autonomous navigation processing unit (130).
[0079] In one embodiment, generating the transformed image (502) may include detecting one or more objects contained in the original image (501).
[0080] In the present disclosure, the original image (501) may mean an image collected by a camera. The original image (501) may also mean an image on which simple preprocessing has been performed on an image collected by a camera.
[0081] In the present disclosure, the original image (501) may include one or more objects. For example, the original image (501) may include any of a variety of objects, such as a ship, a tube, a buoy, a lighthouse, trash, a drowning person, a glacier, or an animal. In one embodiment, the learning image generation unit (500) may detect one or more objects included in the original image (501).
[0082] In one embodiment, a transformed image (502) may be generated based on one or more detected objects. The learning image generation unit (500) may sequentially perform two major processes to achieve natural synthesis when moving a detected object in an image or replacing or adding an object retrieved from an object library to an existing image.
[0083] First, the learning image generation unit (500) can normalize the color characteristics of the object based on at least one of the average color, brightness, and contrast of the background so that the object has similar illuminance and color sense to the background atmosphere. Thereafter, the learning image generation unit (500) can adjust the mixing ratio with the background pixels according to the distance from the center of the object to form a natural boundary, so that the pixel values are mixed in such a way that the closer the object is to the boundary of the synthesized object, the more the background pixels are reflected. As a result, the visual incongruity between the background and the synthesized object is alleviated, enabling more realistic synthesis.
[0084] In one embodiment, the transformed image (502) can be any one of a move-transform image, a replace-transform image, and an add-transform image.
[0085] In the present disclosure, a movement-deformation image may be an image generated by a process of moving at least one of one or more objects included in an original image (501). That is, the learning image generation unit (500) may generate a deformed image (502) by moving at least one of one or more objects included in an original image (501).
[0086] In the present disclosure, a replacement-deformation image may be an image generated by a process of replacing at least one of one or more objects included in an original image (501). That is, the learning image generation unit (500) may generate a deformed image (502) by replacing at least one of one or more objects included in an original image (501).
[0087] In one embodiment, replacing at least one of the one or more objects may involve replacing at least one of the one or more objects with an object of the same type but of a different appearance, or with an object of a different type. For example, if the original image (501) includes a ship, the ship may be replaced with a ship of a different appearance. For example, if the original image (501) includes a ship, the ship may be replaced with a buoy, which is a different type of object.
[0088] In one embodiment, the additive-deformed image may be an image generated by a process of adding an object other than one or more objects included in the original image (501). That is, the learning image generation unit (500) may generate a modified image (502) by adding an object other than one or more objects included in the original image (501).
[0089] In one embodiment, the transformed image (502) may be an image based on a combination of two or more of a move-transform image, a replace-transform image, and an add-transform image. That is, the transformed image (502) may be generated based on a combination of two or more of a process of moving at least one of one or more objects, a process of replacing at least one of one or more objects, and a process of adding one or more objects to another object. Since marine objects float on the water surface, reflections on the water inevitably occur. Therefore, in order to reflect visual elements unique to the marine environment during object synthesis, the learning image generation unit (500) may perform the following process. The learning image generation unit (500) may flip the object to be synthesized around the vertical axis to generate a mirror image. In addition, the learning image generation unit (500) may apply a Gaussian blur and a vertical blending technique to the object to be synthesized. The Gaussian blur is a technique applied in consideration of the water reflection characteristic that makes the actual object appear fainter than the actual object. In particular, in order to reflect the characteristic that the water surface reflection becomes fainter as the image goes downward, a weighted average between the mirror image of the object and the surrounding background image can be calculated. The learning image generation unit (500) can mix pixel values in a way that the surrounding background pixels are reflected more as the vertical distance of each pixel center from the water surface boundary increases. In addition, the learning image generation unit (500) can apply a sine wave-based coordinate transformation technique to express the water surface distortion of an object replaced or added to the image.
[0090] In one embodiment, the learning image generation unit (500) may generate a modified image (502) by referencing an object library. The object library may be understood as a database in which information about objects is stored. The object library may include information about objects required in the process of generating the modified image (502). For example, the object library may include information about the types of one or more objects and the appearance of one or more objects corresponding to each of the types of one or more objects. The learning image generation unit (500) may retrieve a specific object to be synthesized from the object library, and before synthesizing it into a modified image, provide a pre-defined prompt as input to a text-based image generation model to reflect various maritime conditions, thereby generating an object with an appearance that meets the conditions. At this time, the generated object includes visual characteristics that reflect marine environmental conditions such as lighting conditions (light scattering, backlighting), weather conditions (sea fog), and a modified appearance (partial submergence), and such visual characteristics may contribute to generating learning data suitable for actual maritime navigation situations. As an example, a predefined prompt may consist of [object name], [lighting condition], [weather condition], and [modified appearance]. Depending on the embodiment, some of [lighting condition], [weather condition], and [modified appearance] may be omitted. An example of a predefined prompt expressed as a comma-separated sentence is "A boat at sunset, in dense sea fog, appearing as a dark silhouette."
[0091] For example, when the learning image generation unit (500) replaces an object included in the original image (501) with an object of the same type but of a different appearance as the object, the learning image generation unit (500) can refer to an object library to determine one of the appearances of one or more objects of the same type as the type of object detected from the original image (501), and generate a modified image (502). For example, when the learning image generation unit (500) replaces a ship included in the original image (501) with a ship of a different appearance, the learning image generation unit (500) can generate a modified image (502) based on one of the appearances of one or more ships stored in the object library.
[0092] For example, when the learning image generation unit (500) replaces an object included in the original image (501) with an object of a different type from the object, the learning image generation unit (500) can refer to the object library to determine the type of the object detected from the original image (501) and one of the different types, and can determine one of the appearances of one or more objects of the determined type to generate a modified image (502). For example, when the learning image generation unit (500) replaces a ship included in the original image (501) with an object of a different type, for example, a buoy, the learning image generation unit (500) can determine one of the buoys of a different type from the ship stored in the object library, and can generate a modified image (502) based on one of the appearances of one or more buoys stored in the object library.
[0093] In one embodiment, the learning image generation unit (500) can update the library based on information about one or more objects included in the original image (501). That is, as described above, the learning image generation unit (500) can detect one or more objects and update the library by storing information about the one or more detected objects in the library. For example, if the learning image generation unit (500) detects a ship in the original image (501), the learning image generation unit (500) can add the appearance of the detected ship to the information about the appearance of the ship in the library.
[0094] In one embodiment, the learning image generation unit (500) can generate a deformed image (502) in a stepwise manner. In one embodiment, the learning image generation unit (500) can generate the deformed image (502) by referring to a deformation manual. The deformation manual can be preset to include information about the deformation steps, including the order of generating the deformed image (502), i.e., the method of deforming the original image (501) and the order of applying the method. Based on the deformation manual, the learning image generation unit (500) can generate the deformed image (502) according to a series of orders, even without a user's setting.
[0095] For example, the transformation manual may include information about transformation steps consisting of transformation steps 1 to 4. For example, transformation step 1 may be a step of moving a first object included in the original image (501) to transform the original image (501), transformation step 2 may be a step of replacing the first object included in the original image (501) with another object of the same type and shape as the first object, transformation step 3 may be a step of replacing the first object included in the original image (501) with an arbitrary object different from the first object, and transformation step 4 may be a step of adding an object different from the first object included in the original image (501). Here, the first object may be the largest object among one or more objects included in the original image (501), specifically, the object occupying the largest area. According to the present example, at least four transformed images (502) may be generated based on the transformation manual. In addition to the present example, the transformation manual may be configured in any form. In one embodiment, the learning image generation unit (500) can generate a transformed image (502) based on a transformation manual if there is no separate setting by the user.
[0096] Figure 6 is a conceptual diagram for explaining the concept of setting data according to one embodiment of the present disclosure.
[0097] The learning image generation unit (600) and the original image (601) illustrated in FIG. 6 may correspond to the learning image generation unit (500) and the original image (501) illustrated in FIG. 5, respectively.
[0098] In one embodiment, the learning image generation unit (600) may receive deformation type setting data (602). The deformation type setting data (602) may be input by a user. The user may input the deformation type setting data (602) through a vessel or a user device for accessing the vessel.
[0099] In the present disclosure, the deformation type setting data (602) is data for setting the type of deformation, and specifically, may be data for determining which image to generate among each of a movement-deformation image, a replacement-deformation image, and an addition-deformation image, or a combination of two or more of a movement-deformation image, a replacement-deformation image, and an addition-deformation image. The deformation type setting data (602) may be set in consideration of the learning purpose, learning direction, vulnerabilities, etc. of the object detection model.
[0100] In one embodiment, the learning image generation unit (600) can transform the original image (601) based on the transformation type setting data (602).
[0101] In one embodiment, the learning image generation unit (600) may receive object type setting data (603). The object type setting data (603) may be input by a user. The user may input the object type setting data (603) through a vessel or a user device for accessing the vessel.
[0102] In the present disclosure, the object type setting data (603) is data for setting the type of an object when transformed, and may be data for an object to be replaced or added during the process of transforming the original image (601). In particular, when generating a replacement-transformation image and an addition-transformation image, or an image in which a replacement-transformation image or an addition-transformation image is combined, the object to be replaced or the object to be added may be determined based on the object type setting data (603). The object type setting data (603) may be set in consideration of the learning purpose, learning direction, vulnerabilities, etc. of the object detection model.
[0103] As described above, the learning image generation unit (600) can generate a transformed image (502) based on a transformation manual if there is no separate setting by the user. Here, the separate setting by the user may include transformation type setting data (602) and object type setting data (603).
[0104] In one embodiment, the learning image generation unit (600) may include a configuration for performing operations for processing an original image according to the various embodiments described above. Specifically, the learning image generation unit (600) may include a model for detecting an object included in an original image, a model for separating a detected object from the original image, a model for moving a detected object or replacing a detected object with another object, and / or a model for adding another object to the original image.
[0105] FIGS. 7A to 7E are drawings illustrating examples of deformed images according to embodiments of the present disclosure.
[0106] Figure 7a is a drawing showing an example of an original image.
[0107] Referring to FIG. 7A, a ship (701), which is an object included in an image (710), is illustrated. The image (710) illustrated in FIG. 7A may be an original image. The learning image generation unit (600) or the autonomous navigation processing unit (130) may detect the ship (701).
[0108] Figure 7b is a drawing showing an example of a movement-deformation image.
[0109] As described above, a translation-deformation image may be an image generated by a process of translating at least one of one or more objects included in an original image.
[0110] Referring to FIG. 7b, an image (710) including a vessel (702) created by moving the vessel (701) of FIG. 7a is shown.
[0111] Figure 7c is a drawing showing an example of a replacement-deformation image.
[0112] As described above, the replacement-transformation image may be an image generated by a process of replacing at least one of the objects included in the original image (501). Furthermore, replacing at least one of the objects may be replacing at least one of the objects with an object of a different type.
[0113] Referring to FIG. 7c, an image (710) is shown in which the vessel (701) of FIG. 7a is replaced with a buoy (703).
[0114] Figure 7d is a drawing showing an example of a replacement-deformation image.
[0115] As described above, replacing at least one of the one or more objects may be replacing at least one of the one or more objects with an object of the same type and of a different appearance.
[0116] Referring to FIG. 7d, an image (710) is shown in which the vessel (701) of FIG. 7a is replaced with a vessel (704) of a different appearance.
[0117] Figure 7e is a drawing showing an example of an additive-deformed image.
[0118] As described above, an additive-modified image may be an image created by a process of adding one or more objects contained in an original image to another object.
[0119] Referring to FIG. 7e, an image (710) is created by adding a person (705) who fell into the water to the ship (701) of FIG. 7a.
[0120] FIG. 8 is a flowchart illustrating a method for generating a deformed image according to another embodiment of the present disclosure.
[0121] In one embodiment, the learning image generation unit (500) may generate a modified image based on an image stored in a database. The database may include base images generated by the learning image generation unit (500). Generating a modified image based on a database has the advantage of generating a larger amount of learning data.
[0122] Referring to FIG. 8, in step 801, the learning image generation unit (500) can generate a first base image and store the first base image in a database.
[0123] In the present disclosure, a base image may refer to an image that serves as a basis for generating a transformed image. In one embodiment, the learning image generation unit (500) may generate a first base image based on a received image, wherein the received image may be an original image collected by a camera.
[0124] In one embodiment, the base image may be generated by deleting at least a portion of one or more objects included in the original image. For example, the base image may be generated by deleting the ship (701) from the original image illustrated in FIG. 7A.
[0125] In one embodiment, the learning image generation unit (500) can store the generated base image in a database, whereby one or more base images can be accumulated and stored in the database.
[0126] Referring to FIG. 8, in step 802, the learning image generation unit (500) can load a second base image from the database.
[0127] Here, the second base image is any base image among one or more base images stored in the database, and may be the same as or different from the first base image.
[0128] Referring to FIG. 8, in step 803, the learning image generation unit (500) can add a third object to the second base image.
[0129] Here, the third object may refer to any object. For example, the third object may be randomly selected from among multiple objects stored in the aforementioned object library.
[0130] In one embodiment, the base image may include information regarding the location where the third object is to be added. The learning image generation unit (500) may add the third object based on the information regarding the location where the third object is to be added.
[0131] Meanwhile, as described above, the learning image generation unit (500) can generate labeling data (503) together with the transformed image (502).
[0132] The learning image generation unit (500) can generate labeling data (503) based on information (e.g., an object library) about one or more objects, replaced objects, or added objects detected from the original image.
[0133] In the present disclosure, the learning image generation unit (600) may include a configuration for performing operations for processing an original image according to the various embodiments described above. Specifically, the learning image generation unit (600) may include a model for detecting an object included in an original image, a model for separating a detected object from the original image, a model for deleting a detected object, a model for moving a detected object or replacing a detected object with another object, and / or a model for adding another object to the original image.
[0134] As described above, the object detection model (510) can be trained based on the transformed image (502) and labeling data (503). The autonomous navigation processing device (130) can train the object detection model (510) based on the transformed image (502) and labeling data (503).
[0135] In the present disclosure, the autonomous navigation processing device (130) can detect objects existing around the ship based on an object detection model (510). Specifically, the autonomous navigation processing device (130) can input images collected through a camera into the object detection model (510) and receive information about objects output by the object detection model (510).
[0136] In the present disclosure, the autonomous navigation processing device (130) can generate or modify a path plan or a docking plan and generate a map based on detected objects. The autonomous navigation processing device (130) can calculate control values for controlling the movement of the vessel based on the generated path plan or docking plan. The autonomous navigation processing device (130) can control the vessel based on the calculated control values.
[0137] In the present disclosure, the autonomous navigation processing device (130) can generate an information provision interface based on the detected object. The information provision interface can be configured to display any information related to the vessel, such as information about the detected object, information about the route plan, information about the generated map, and information about the automatic docking control process.
[0138] In the present disclosure, the autonomous navigation processing device (130) may provide a generated information provision interface to a user. For example, the autonomous navigation processing device (130) may display the generated information provision interface via a display device communicatively connected to the autonomous navigation processing device (130). The display device may be mounted on the vessel or included in a user device for accessing the vessel.
[0139] FIG. 9 is a block diagram of a device according to one embodiment of the present disclosure.
[0140] The device (900) illustrated in FIG. 9 is an object detection device, and the object detection model device may be the autonomous operation processing device (130) described above or may be included in the autonomous operation processing device (130).
[0141] Referring to FIG. 9, the device (900) may include a communication unit (910), a processor (920), and a database (930). Only components related to the embodiment are illustrated in the device (900) of FIG. 9. Therefore, those skilled in the art will appreciate that other general components may be included in addition to the components illustrated in FIG. 9.
[0142] The communication unit (910) may include one or more components that enable wired / wireless communication with an external server or external device. For example, the communication unit (910) may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast receiving unit (not shown).
[0143] DB (930) is hardware that stores various data processed within the device (900), and can store programs for processing and controlling the processor (920). DB (930) can store payment information, user information, etc.
[0144] DB (930) may include random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disk storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.
[0145] The processor (920) controls the overall operation of the device (900). For example, the processor (920) can control the input unit (not shown), the display (not shown), the communication unit (910), the DB (930), etc., by executing programs stored in the DB (930). The processor (920) can control the operation of the device (900) by executing programs stored in the DB (930).
[0146] The processor (920) can control at least some of the operations of the device (900) described above in FIGS. 1 to 8.
[0147] The processor (920) may be implemented using at least one of application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.
[0148] In one embodiment, the device (900) may be a mobile electronic device. For example, the device (900) may be implemented as a smartphone, tablet PC, PC, smart TV, personal digital assistant (PDA), laptop, media player, navigation device, camera-equipped device, or other mobile electronic device. Furthermore, the device (900) may be implemented as a wearable device, such as a watch, glasses, hair band, or ring, equipped with communication and data processing capabilities.
[0149] Embodiments according to the present invention may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions, such as ROMs, RAMs, and flash memories.
[0150] Meanwhile, the computer program may be specifically designed and constructed for the present invention, or may be one known and available to those skilled in the computer software field. Examples of computer programs may include not only machine language code, such as that generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
[0151] According to one embodiment, the method according to various embodiments of the present disclosure may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0152] Unless the steps constituting the method according to the present invention are explicitly described in a specific order or are otherwise described in a different order, the steps may be performed in any appropriate order. The present invention is not necessarily limited to the order in which the steps are described. The use of all examples or exemplary terms (e.g., “for example,” “etc.”) in the present invention is merely intended to illustrate the present invention in more detail, and the scope of the present invention is not limited by the examples or exemplary terms unless otherwise defined by the claims. Furthermore, those skilled in the art will appreciate that various modifications, combinations, and variations can be configured according to design conditions and factors within the scope of the appended claims or their equivalents.
[0153] Therefore, the idea of the present invention should not be limited to the embodiments described above, and not only the scope of the patent claims described below but also all scopes equivalent to or equivalently modified from the scope of the patent claims are considered to fall within the scope of the idea of the present invention.
Claims
1. As an object detection method for autonomous driving, A step of receiving an original image collected by a camera; A step of generating a transformed image by transforming the original image; For learning an object detection model, a step of inputting the transformed image into the object detection model; and A step of detecting an object existing around a ship based on the above object detection model; including, method.
2. In paragraph 1, The step of generating the above-described transformed image is: A step of detecting one or more objects included in the original image; Including, The above modified image is, generated based on one or more of the above objects, method.
3. In paragraph 2, The above modified image is, generated based on at least one of the steps of moving at least one of the above one or more objects, replacing at least one of the above one or more objects, or adding another object to the above one or more objects. method.
4. In paragraph 1, A step of creating a base image by deleting at least some of one or more objects included in the original image and storing the base image in a database; including more, method.
5. In paragraph 1, The step of generating the above-described transformed image is: a step of loading a base image from the above database; and A step of adding a third object to the above base image; including, method.
6. In paragraph 1, The step of generating the above-described transformed image is: A step of generating labeling data for each of one or more objects included in the transformed image; including, method.
7. In paragraph 1, A step of receiving transformation type setting data and object type setting data input by a user; Including more, The creation of the above transformed image is as follows: Based on the above transformation type setting data and the above object type setting data, method.
8. In paragraph 1, A step of generating an information providing interface based on the detected object; including more, method.
9. In paragraph 2, The above modified image is, A color characteristic generated based on one or more objects normalized based on at least one of the average color, brightness, and contrast of the background, method.
10. In paragraph 9, The above modified image is, By adjusting the blending ratio with the background pixels according to the distance from the center of the one or more objects, the closer to the boundary of the one or more objects, the more background pixels are reflected and blended, and generated. method.
11. In paragraph 2, The above modified image is, An image that generates a mirror image of one or more objects and represents an object reflected on a surface of water based on the mirror image, method.
12. In paragraph 2, The above modified image is, It is generated based on the result of applying at least one of a Gaussian blur technique and a sine wave-based coordinate transformation technique to one or more of the objects. method.
13. In paragraph 1, The step of generating a transformed image by transforming the original image is as follows: After loading an object from the object library, change at least one of the surface condition, weather condition and transformed appearance of the object according to a pre-defined prompt, The above modified image is, Created based on the above changed object, method.
14. As an object detection device for autonomous operation, As a ship device, memory in which at least one program is stored; and A processor that operates by executing at least one program; The above processor, Receive the original image collected by the camera, Generate a transformed image by transforming the original image above, To train the object detection model, the transformed image is input into the object detection model, Based on the above object detection model, detecting objects existing around the ship, device.
15. A computer-readable recording medium recording a program for executing the method according to paragraph 1 on a computer.
Citation Information
Patent Citations
Program, information storage medium, and image forming system
JP2006318388A
Composition for epoxy-based structural adhesive
KR1020230067318A
Smart ship system and operation method thereof
KR1020230174591A
Robot and method of controlling the same
KR1020250074407A
Multi data set-based learning apparatus and method for multi object recognition machine learning
KR102416691B1