Method and apparatus for generating improved training data images

The method improves object recognition in autonomous vehicles by integrating multiple recognition techniques and active learning to generate enhanced training data, addressing the limitations of existing systems and enhancing accuracy and efficiency.

JP7756184B2Active Publication Date: 2025-10-1742DOT INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024022949
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-21
Filing Date
2024-02-19
Publication Date
2025-10-17
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

Existing methods for object recognition and distance estimation in autonomous vehicles face challenges due to the variability of road objects, limitations in processing capabilities, and inaccuracies in camera-based and radar-based systems, with LiDAR being costly and environmentally sensitive.

Method used

A method involving multiple recognition and detection techniques to generate improved training data images by integrating and sampling frame sets, using algorithms like YoloV4-CSP and YoloV4-P7, and correcting object recognition results through coordinate transformations and active learning.

Benefits of technology

Enhances the object recognition rate of autonomous vehicles by generating more effective training data, improving the accuracy and efficiency of object detection and distance estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756184000013
    Figure 0007756184000013
  • Figure 0007756184000014
    Figure 0007756184000014
  • Figure 0007756184000015
    Figure 0007756184000015
Patent Text Reader

Abstract

To provide a method and device for generating an improved learning data video.SOLUTION: A method includes a step S1410 for applying at least two or more recognition techniques to a first video acquired during travel to recognize an object included in the first video, a step S1430 for applying at least two or more detection techniques to a result of recognizing the object to detect frames by applied detection techniques, a step S1450 for integrating detected frames to generate a frame set including a plurality of frames, and a step S1470 for sampling the integrated frame set to generate a second video.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for generating training data images, and more particularly to an improved method for generating training data for improving the ability of an autonomous vehicle capable of recognizing objects and operating autonomously to recognize objects on a road while traveling, and an apparatus for implementing the method. [Background technology]

[0002] The convergence of information and communications technology and the automotive industry is rapidly making vehicles smarter. This smartification is driving vehicles from simple mechanical devices to smart cars, with autonomous driving attracting particular attention as a core technology for smart cars. Autonomous driving is a technology in which an autonomous driving module installed in a vehicle actively controls the vehicle's driving state, allowing the vehicle to reach its destination on its own, without the driver having to operate the steering wheel, accelerator pedal, brake, etc.

[0003] In order to enable autonomous vehicles to drive safely, various research efforts have been conducted on methods for vehicles to accurately recognize pedestrians and other vehicles during autonomous driving and calculate the distance to recognized objects. However, the characteristics of objects that can appear on the road while the vehicle is driving are virtually infinite, and there are limitations to the processing capabilities of the modules installed in autonomous vehicles, so there is currently no known method for perfectly recognizing objects on the road.

[0004] In camera-based object recognition and distance estimation, objects in the real 3D world are projected onto a 2D image, resulting in significant loss of distance information. In particular, the large variability in features often used to calculate pedestrian position (pedestrian height and ground contact points) leads to significant errors.

[0005] In the case of object recognition and distance estimation using radar, due to the characteristics of the radio waves used by radar, the ability to quickly identify and classify objects is poor, making it difficult to determine whether an object is a pedestrian or a vehicle. In particular, in the case of pedestrians and two-wheeled vehicles (bicycles and motorcycles) on the road, the recognition results tend to be even worse due to weak signal strength.

[0006] In recent years, object recognition and distance estimation technology using LiDAR has been attracting attention due to its relatively high accuracy. However, because high-power lasers are dangerous, LiDAR must operate based on reduced-power lasers. Unlike the radio waves used by radar, lasers are greatly affected by the surrounding environment. Furthermore, the high cost of LiDAR sensors has been pointed out as limitations.

[0007] The above-mentioned background art is technical information that the inventor possessed in order to derive the present invention or that he acquired in the process of deriving the present invention, and is not necessarily publicly known art that was disclosed to the general public prior to the filing of the present invention. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Korean Patent Registration No. 10-2438114 (August 25, 2022) Summary of the Invention [Problem to be solved by the invention]

[0009] The technical problem to be solved by the present invention is to provide a method for generating improved training data images. [Means for solving the problem]

[0010] A method according to one embodiment of the present invention for solving the above technical problem includes the steps of applying at least two or more recognition techniques to a first image acquired while driving to recognize objects included in the first image, applying at least two or more detection techniques to the results of the object recognition to detect frames for each of the applied detection techniques, integrating the detected frames to generate a frame set including a plurality of frames, and sampling the integrated frame set to generate a second image.

[0011] In the method, the step of generating the second image includes a step of generating a frame group including at least one frame based on the integrated frame set, the frame group including non-overlapping frames, and a step of extracting frames for each frame group to generate the second image.

[0012] In the method, generating the second image includes extracting one frame from each of the frame groups to generate the second image.

[0013] In the method, generating the second image includes extracting, for each frame group, a number of frames corresponding to a weight set for each frame group to generate the second image.

[0014] In the method, the weight set for each frame group may be a value determined based on the number of frames included in each frame group.

[0015] In the above method, the step of generating the second image can include sampling frames included in the integrated frame set based on a predetermined time interval to extract multiple frames, and generating the second image using the extracted frames.

[0016] In the method, in the step of generating the integrated frame set, frames detected as overlapping among frames detected for each detection technique may be identified, and in the step of generating the second image, the second image may be generated by essentially including the frames detected as overlapping.

[0017] In the method, in the step of recognizing an object included in the first image, a first recognition technique and a second recognition technique may be applied to recognize the object included in the first image, and the first recognition technique may be an algorithm for recognizing the object in the first image based on YoloV4-CSP, and the second recognition technique may be an algorithm for recognizing the object in the first image based on YoloV4-P7.

[0018] In the method, the at least two or more recognition techniques may include a first recognition technique and a second recognition technique, and the at least two or more detection techniques may include a detection technique that detects a frame based on a result of comparing frames of objects recognized by the first recognition technique and the second recognition technique, respectively.

[0019] In the method, the at least two or more detection techniques may include a detection technique that detects a frame based on the result of detecting that an object recognized from the first image has disappeared for a predetermined period of time and then reappeared.

[0020] According to another embodiment of the present invention, an apparatus for solving the above technical problems includes a memory storing at least one program, and a processor that performs calculations by executing the at least one program, wherein the processor applies at least two or more recognition techniques to a first image acquired while driving to recognize objects included in the first image, applies at least two or more detection techniques to the results of the object recognition to detect frames for each of the applied detection techniques, integrates the detected frames to generate a frame set including a plurality of frames, and samples the integrated frame set to generate a second image.

[0021] A method according to one embodiment of the present invention for solving the above technical problem includes the steps of recognizing an object from an image, identifying a first frame in which the object is recognized and a second frame in which the object is not recognized, generating a first outline of the object in the first frame and obtaining coordinates of a first cuboid of the object based on first coordinate values ​​constituting the first outline, determining whether the object should be recognized in the second frame and generating a second outline of the object in the second frame based on the determination result, calculating a conversion relationship value between the first coordinate values ​​constituting the first outline and second coordinate values ​​constituting the second outline, and calculating coordinates of a second cuboid of the object in the second frame by applying the conversion relationship value to the coordinates of the first cuboid.

[0022] According to another embodiment of the present invention, an apparatus for solving the above technical problem includes a memory having at least one program stored therein, and a processor that performs calculations by executing the at least one program, wherein the processor recognizes an object from an image, identifies a first frame in which the object is recognized and a second frame in which the object is not recognized, generates a first outline of the object in the first frame, obtains coordinates of a first cuboid of the object based on first coordinate values ​​constituting the first outline, determines whether the object should be recognized in the second frame, generates a second outline of the object in the second frame based on the result of the determination, calculates a conversion relationship value between the first coordinate values ​​constituting the first outline and the second coordinate values ​​constituting the second outline, and applies the conversion relationship value to the coordinates of the first cuboid to calculate the coordinates of the second cuboid of the object in the second frame.

[0023] A method according to one embodiment of the present invention for solving the above technical problems includes the steps of recognizing an object included in an image using a first recognition algorithm that recognizes an object for each frame of the image, forming a track using a plurality of frames included in the image and recognizing the object using a second recognition algorithm that recognizes the object included in the track, comparing a result of recognizing the object using the first recognition algorithm with a result of recognizing the object using the second recognition algorithm, and correcting the results of recognizing the object in the image using the first recognition algorithm and the second recognition algorithm based on the result of the comparison.

[0024] According to another embodiment of the present invention, an apparatus for solving the above technical problems includes a memory having at least one program stored therein, and a processor that performs calculations by executing the at least one program, wherein the processor recognizes an object included in an image using a first recognition algorithm that recognizes an object for each frame of the image, forms a track using a plurality of frames included in the image, recognizes the object using a second recognition algorithm that recognizes the object included in the track, compares a result of recognizing the object using the first recognition algorithm with a result of recognizing the object using the second recognition algorithm, and corrects the results of recognizing the object in the image using the first recognition algorithm and the second recognition algorithm based on the result of the comparison.

[0025] An embodiment of the present invention can provide a computer-readable recording medium storing a program for executing the method. [Effects of the Invention]

[0026] According to the present invention, it is possible to obtain learning data that can improve the object recognition rate of an autonomous vehicle by capturing images of the vehicle while it is traveling using a camera attached to the autonomous vehicle.

[0027] In particular, the training data generated by the present invention can more efficiently train the object recognition rate of an object recognition device for an autonomous vehicle compared to existing training data. [Brief explanation of the drawings]

[0028] [Figure 1] FIG. 1 is a diagram illustrating an autonomous driving method according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an autonomous driving method according to an embodiment. [Figure 3] FIG. 1 is a diagram illustrating an autonomous driving method according to an embodiment. [Figure 4A]FIG. 1 is a diagram of a camera capturing an image of the exterior of a vehicle according to one embodiment. [Figure 4B] FIG. 1 is a diagram of a camera capturing an image of the exterior of a vehicle according to one embodiment. [Figure 5] FIG. 1 is a schematic diagram illustrating an object recognition method according to an embodiment. [Figure 6] FIG. 1 is a diagram conceptually illustrating a method for improving object recognition rates in an autonomous vehicle, according to an embodiment of the present invention. [Figure 7A] 4 is a diagram illustrating a filtering process performed in an object recognition rate improving apparatus according to an embodiment of the present invention. [Figure 7B] 4 is a diagram illustrating a filtering process performed in an object recognition rate improving apparatus according to an embodiment of the present invention. [Figure 7C] 4 is a diagram illustrating a filtering process performed in an object recognition rate improving apparatus according to an embodiment of the present invention. [Figure 8] 10 is a diagram illustrating a process of applying active learning to improve the object recognition rate of an autonomous vehicle according to another embodiment of the present invention. FIG. [Figure 9] 1 is a flowchart illustrating an example of a method for improving an object recognition rate according to the present invention. [Figure 10] 10A and 10B are diagrams illustrating a method for improving an object recognition rate according to still another embodiment of the present invention. [Figure 11] 11 is a flowchart illustrating a method for improving an object recognition rate according to the embodiment described in FIG. 10. [Figure 12] FIG. 1 is a conceptual diagram conceptually illustrating an example of a method for generating improved training data videos according to the present invention. [Figure 13] FIG. 13 is a diagrammatic illustration of an expanded concept of the embodiment described in FIG. 12. [Figure 14] 4 is a flowchart illustrating an example of a second video generating method according to the present invention. [Figure 15] FIG. 2 is a block diagram of a second image generating device according to an embodiment. [Figure 16] 11 is a flowchart illustrating a method for improving an object recognition rate according to the embodiment described in FIG. 10. [Figure 17] 1 is a diagram for explaining cuboids of objects in an image according to the present invention; [Figure 18] 1 is a schematic diagram for explaining a method for obtaining cuboids of an object in an image according to the present invention; [Figure 19] 10A and 10B are diagrams for explaining a linear mapping method performed by the cuboid acquisition device according to the present invention. [Figure 20] 10A and 10B are diagrams for explaining another example of a linear mapping method performed by the cuboid acquisition device according to the present invention. [Figure 21] 1 is a flowchart showing an example of a cuboid obtaining method according to the present invention. [Figure 22] FIG. 1 is a block diagram of a cuboid acquisition device according to one embodiment. [Figure 23] 1 is a schematic diagram for explaining a method for improving an object recognition rate according to the present invention; [Figure 24] FIG. 24 is a diagram for explaining the processing process after the track registration explained in step S2380 of FIG. [Figure 25] FIG. 24 is a diagram for explaining the processing process after the track deletion described in step S2360 of FIG. [Figure 26] 1 is a flowchart illustrating an example of a method for improving object recognition rate according to the present invention. [Figure 27] 1 is a block diagram of an object recognition rate improving apparatus according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0029] The present invention can be modified in various ways and has various embodiments, so specific embodiments are illustrated in the drawings and described in detail in the detailed description. The advantages and features of the present invention, as well as methods for achieving them, will become apparent by referring to the embodiments described in detail below in conjunction with the drawings. However, the present invention is not limited to the embodiments disclosed below, and can be realized in various forms.

[0030] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. In the description with reference to the drawings, the same or corresponding components will be given the same reference numerals, and duplicate descriptions thereof will be omitted.

[0031] In the following embodiments, terms such as first and second are not used in a limiting sense but are used to distinguish one component from another.

[0032] In the following embodiments, singular expressions include plural expressions unless the context clearly indicates otherwise.

[0033] In the following embodiments, terms such as "comprise" and "have" mean that the features or components described in the specification are present, and do not preclude the possibility that one or more other features or components may be added.

[0034] In some embodiments, the order of certain steps may be different from that described, for example, two steps described as successive may be performed substantially simultaneously or in the reverse order from that described.

[0035] 1 to 3 are diagrams for explaining an autonomous driving method according to an embodiment.

[0036] Referring to FIG. 1 , an autonomous driving device according to an embodiment of the present invention can be mounted on a vehicle to realize an autonomous vehicle 10. The autonomous driving device mounted on the autonomous vehicle 10 may include various sensors for collecting information about surrounding conditions. For example, the autonomous driving device may detect the movement of a leading vehicle 20 traveling ahead using an image sensor and / or an event sensor mounted on the front of the autonomous vehicle 10. The autonomous driving device may further include sensors for detecting other vehicles 30 traveling on adjacent lanes as well as in front of the autonomous vehicle 10, pedestrians around the autonomous vehicle 10, and the like.

[0037] At least one of the sensors for collecting situational information about the autonomous vehicle may have a predetermined field of view (FoV), as shown in Fig. 1. As an example, if a sensor mounted on the front of the autonomous vehicle 10 has a field of view (FoV) as shown in Fig. 1, information detected at the center of the sensor may be of relatively high importance. This is because the information detected at the center of the sensor contains most of the information corresponding to the movement of the leading vehicle 20.

[0038] The autonomous driving device processes information collected by sensors of the autonomous vehicle 10 in real time to control the movement of the autonomous vehicle 10, while at least a portion of the information collected by the sensors can be stored in a memory device.

[0039] 2, the autonomous driving device 40 may include a sensor unit 41, a processor 46, a memory system 47, a vehicle control module 48, etc. The sensor unit 41 includes a plurality of sensors 42 to 45, which may include an image sensor, an event sensor, an illuminance sensor, a GPS device, an acceleration sensor, etc.

[0040] The data collected by the sensors 42 to 45 may be transmitted to the processor 46. The processor 46 stores the data collected by the sensors 42 to 45 in a memory system 47 and controls a body control module 48 to determine the movement of the vehicle based on the data collected by the sensors 42 to 45. The memory system 47 may include two or more memory devices and a system controller for controlling the memory devices. Each of the memory devices may be provided as a single semiconductor chip.

[0041] In addition to the system controller of memory system 47, each of the memory devices included in memory system 47 may include a memory controller, and the memory controller may include an artificial intelligence (AI) calculation circuit such as a neural network. The memory controller can generate calculation data by applying a predetermined weight to data received from sensors 42 to 45 or processor 46, and store the calculation data in a memory chip.

[0042] 3 is a diagram showing an example of video data acquired by a sensor of an autonomous vehicle equipped with an autonomous driving device. Referring to FIG. 3, video data 50 may be data acquired by a sensor attached to the front of the autonomous vehicle. Therefore, video data 50 may include a front portion 51 of the autonomous vehicle, a preceding vehicle 52 on the same road as the autonomous vehicle, vehicles 53 traveling around the autonomous vehicle, and an area of ​​no interest 54.

[0043] 3, the data of the area in which the front portion 51 of the autonomous vehicle and the area of ​​no interest 54 are displayed may be data that is unlikely to affect the operation of the autonomous vehicle. In other words, the front portion 51 of the autonomous vehicle and the area of ​​no interest 54 may be considered to be data with a relatively low level of importance.

[0044] On the other hand, the distance to the preceding vehicle 52, the lane-changing behavior of the moving vehicle 53, etc. can be very important factors in the safe operation of an autonomous vehicle. Therefore, in the video data 50, data of the area including the preceding vehicle 52 and the moving vehicle 53 can be relatively important in the operation of the autonomous vehicle.

[0045] The memory device of the autonomous driving device can store the image data 50 received from the sensor by assigning different weights to each area. For example, a high weight can be assigned to data of an area including a leading vehicle 52 and a moving vehicle 53, and a low weight can be assigned to data of an area including the front part 51 of the autonomous driving vehicle and an area of ​​no interest 54.

[0046] 4A and 4B are diagrams of a camera capturing an image of the exterior of a vehicle according to one embodiment.

[0047] The camera is mounted on a vehicle and can capture images of the exterior of the vehicle. The camera can capture images of the front, side, rear, etc. of the vehicle. The object recognition rate improving device according to the present invention can acquire multiple images captured by the camera. The multiple images captured by the camera may include multiple objects.

[0048] Information about an object includes object type information and object attribute information. Here, object type information is index information indicating the type of object and is composed of a broad group and a narrow class. Also, object attribute information indicates attribute information related to the current state of the object and includes information such as movement information, rotation information, traffic information, color information, and visibility information.

[0049] In one embodiment, the groups and classes included in the object type information are as shown in Table 1 below, but are not limited thereto. [Table 1]

[0050] The motion information indicates the motion information of an object and can be defined as stopped, parked, moving, etc. In the case of a vehicle, stopped, parked, or moving can be determined as the object attribute information, and in the case of a pedestrian, moving, stopped, or unknown can be determined as the object attribute information, and in the case of a stationary object such as a traffic light, the default value of stillness can be determined as the object attribute information.

[0051] Rotation information indicates rotation information of an object and can be defined as front, rear, horizontal, vertical, side, etc. In the case of a vehicle, front, rear, or side can be determined as object attribute information, and for a horizontal or vertical traffic light, horizontal or vertical can be determined as object attribute information, respectively.

[0052] Traffic information means traffic information of an object, and can be defined as traffic sign instructions, cautions, regulations, auxiliary signs, etc. Color information means color information of an object, and can indicate the color of the object, the color of a traffic light, and the color of a traffic sign.

[0053] 4A , the object 411 may be a pedestrian. The image 410 may have a predetermined size. The same object 411 may be included in multiple images 410, but the relative position between the vehicle and the object 411 continues to change as the vehicle travels on the road, and the object 411 also moves over time, so that the position of the same object 411 changes in each image.

[0054] Using the entire image to determine what the same object is in each image requires a huge amount of data transmission and computation, making edge computing in vehicle-mounted devices difficult to process and difficult to analyze in real time.

[0055] 4B shows a bounding box 421 included in an image 420. A bounding box is metadata related to an object, and bounding box information may include object type information (group, class, etc.), position information on the image 420, size information, etc.

[0056] Referring to FIG. 4B, the bounding box information may include information that the object 411 belongs to a pedestrian class, information that the top left vertex of the object 411 is located at (x, y) on the image, information that the size of the object 411 is w×h, and current state information (i.e., movement information) that the object 411 is moving.

[0057] FIG. 5 is a schematic diagram illustrating an object recognition method according to one embodiment.

[0058] The object recognition rate improving apparatus may acquire a plurality of frames by separating a video captured by a camera into frames, which may include a previous frame 510 and a current frame 520.

[0059] The object recognition rate improving device can recognize a first pedestrian object 511 in the previous frame 510 .

[0060] In one embodiment, the object recognition rate improving device may divide a frame into grids of equal size, predict the number of bounding boxes specified in a predefined format centered on the center of each grid for each grid, and calculate the confidence level based on the predicted number. The object recognition rate improving device may determine whether an object is included in a frame or whether only the background is present, and select a position having a high object confidence level to determine an object category, thereby ultimately recognizing the object. However, the object recognition method in the present disclosure is not limited thereto.

[0061] The object recognition rate improving device may acquire first position information of the first pedestrian object 511 recognized in the previous frame 510. As described above with reference to FIGS. 4A and 4B, the first position information may include coordinate information and length and width information of any one vertex (e.g., the vertex at the top left corner) of the bounding box corresponding to the first pedestrian object 511 in the previous frame 510.

[0062] In addition, the object recognition rate improving device can acquire second position information of the second pedestrian object 521 recognized in the current frame 520.

[0063] The object recognition rate improvement device can calculate the similarity between the first position information of the first pedestrian object 511 recognized in the previous frame 510 and the second position information of the second pedestrian object 521 recognized in the current frame 520.

[0064] 5, the object recognition rate improving device can use the first position information and the second position information to calculate the intersection and union of the first pedestrian object 511 and the second pedestrian object 521. The object recognition rate improving device calculates the value of the intersection area for the union area, and if the calculated value is equal to or greater than a threshold, can determine that the first pedestrian object 511 and the second pedestrian object 521 are the same pedestrian object.

[0065] However, the method for determining the identity between objects is not limited to the above-mentioned method.

[0066] FIG. 6 is a diagram conceptually illustrating a method for improving object recognition rates in an autonomous vehicle, according to one embodiment of the present invention.

[0067] To summarize one embodiment of the present invention with reference to FIG. 6, in one embodiment of the present invention, when raw data 610 is input into a first model 620 and a second model 630, the result data calculated by each model is received and processed by a variation data calculation module 640 to calculate variation data 645, and the calculated variation data 645 is received and analyzed by a weakness point analysis module 650 to identify weakness points.

[0068] More specifically, in the present invention, raw data 610 refers to video collected by a camera module mounted on an autonomous vehicle. In particular, the raw data 610 is video data that has not been pre-processed after being generated by the camera module, and is composed of multiple frames, and the frame rate may be, but is not limited to, 30 or 60 frames per second.

[0069] The first model 620 refers to a model attached to an autonomous vehicle, which receives the raw data 610 as input data and outputs the results of recognizing objects contained in the raw data 610 as output data.

[0070] The second model 630 is a model included in the server, and similar to the first model 620, receives the raw data 610 as input data and outputs the results of recognizing objects included in the raw data 610 as output data. Compared to the first model 620, which does not have high performance due to limited resources, the second model 630 can be a high-performance model that can use sufficient resources based on a large memory.

[0071] The camera module of the autonomous vehicle is controlled so that the collected raw data 610 is transmitted via the communication module to the first model 620 as well as the second model 630 for processing.

[0072] The output data output from the first model 620 and the second model 630 may include information regarding at least one of the relative positions, sizes, and directions of vehicles, pedestrians, etc. included in each frame of the video.

[0073] In the present invention, the first model 620, being installed in an autonomous vehicle, has relatively limited resources and operates in a limited environment compared to the second model 630. Due to the difference in scale of the models, information regarding the number and types of objects recognized from an image when the raw data 610 is input to the second model 630 may be improved compared to information regarding the number and types of objects recognized from an image when the raw data 610 is input to the first model 620.

[0074] [Table 2] [Table 3] Tables 2 and 3 are examples showing numerically the performance of the first model 620 and the second model 630. More specifically, Table 2 shows the object recognition rate when YoloV4-CSP is used as the first model 620, and Table 3 shows the object recognition rate when YoloV4-P7 is used as the second model 630. Comparing Tables 2 and 3, it can be seen that YoloV4-P7 is generally superior to YoloV4-CSP in terms of the recognition rate of objects included in the raw data 610, such as cars, pedestrians, trucks, buses, two-wheelers, and miscellaneous objects.

[0075] Tables 2 and 3 are illustrative examples of the performance of the first model 620 and the second model 630, and therefore the first model 620 and the second model 630 in the present invention are not limited to YoloV4-CSP and YoloV4-P7 described in Tables 2 and 3, respectively.

[0076] The variability data calculation module 640 may calculate variability data 645 by analyzing output data from the first model 620 and the second model 630. The variability data 645 refers to data relating to the variability between the result of inputting the raw data 610 to the first model 620 and the result of inputting the raw data 610 to the second model 630, and more specifically, may be calculated by comparing the same frames. For example, if the raw data 610 is video data consisting of 10 frames, the variability data 645 may be a result of calculating the variability by comparing the result of inputting the first frame of the raw data 610 to the first model 620 and the result of inputting the first frame of the raw data 610 to the second model 630.

[0077] The variance data calculation module 640 calculates the Intersection over Union (IoU) value of bounding boxes between frames constituting the raw data 610, matches the bounding boxes with the largest IoU, and determines the bounding box detected only in the output data of the second model 630 as a weak point target and transmits it to the weak point analysis module. The method by which the variance data calculation module 640 matches bounding boxes between frames based on the IoU value to calculate variance data has already been described with reference to FIG. 5, so a description thereof will be omitted.

[0078] Hereinafter, the raw data 610 is input to the first model 620 and the output data is referred to as the first recognition result, and the raw data 610 is input to the second model 630 and the output data is referred to as the second recognition result.

[0079] The weak point analysis module 650 receives the variation data from the variation data calculation module 640 and analyzes the weak points. Here, the weak point refers to data related to information about an object that is not detected by the first model 620 but is detected by the second model 630 due to the limited performance of the first model 620, which is installed in the autonomous vehicle and has a relatively smaller amount of calculation than the second model 630. For example, if the second model 630 receives the raw data 610 and recognizes one car and one bus as objects in the image, and the first model 620 receives the raw data 610 and recognizes one car as an object in the image, the weak point may be information about the bus that the first model 620 does not recognize (detect).

[0080] The weak points analyzed by the weak point analysis module 650 can be used as training data to improve the object recognition performance of the first model 620. In addition, the weak points may be pre-processed through a series of pre-processing processes (or filtering processes) to be used as training data for the first model 620, which will be described later.

[0081] 6, the first model 620, the variation data calculation module 640, and the weak point analysis module 650 may be physically or logically included in the device for improving the object recognition rate of an autonomous vehicle according to an embodiment of the present invention. Furthermore, in an actual implementation of the present invention, the first model 620, the second model 630, the variation data calculation module 640, and the weak point analysis module 650 may be called by other names, or may be implemented in a form in which one module is integrated with another.

[0082] 7A to 7C are diagrams illustrating a filtering process performed in an object recognition rate improving apparatus according to an embodiment of the present invention.

[0083] 7A shows the variation data before filtering, and diagrammatically shows that a first object 710a, a second object 720a, a third object 730a, a fourth object 740a, and a fifth object 750a are recognized as objects. More specifically, the five objects shown in FIG. 7A can be understood as not being recognized in the first recognition result but being recognized in the second recognition result, processed as variation data, and transmitted to the weakness point analysis module 650. The weakness point analysis module 650 can filter the variation data according to preset filter criteria to leave only meaningful object information in the variation data.

[0084] For example, the preset filter criterion is a size criterion related to the size of a bounding box included in the variation data, and the weak point analysis module 650 may remove bounding boxes smaller than the size criterion based on the variation data. Here, the size criterion may be a criterion for removing bounding boxes with a height or width of less than 120 pixels. However, since the above values ​​are exemplary, the height or width criterion may vary depending on the embodiment.

[0085] As another example, the preset filter criteria are classification criteria for classifying the object types of bounding boxes included in the variation data, and the weak point analysis module 650 can remove bounding boxes of specific types of objects according to the classification criteria as information based on the variation data. Here, the specific type refers to the class written at the top of the bounding box, and the five bounding boxes in Figure 7A show a total of four classes (passenger car, truck, pedestrian, and motorcycle).

[0086] If the filter criteria set in the weak point analysis module 650 simultaneously include a size criterion for removing bounding boxes with a height of less than 120 pixels or a width of less than 120 pixels, and a classification criterion for removing bounding boxes of pedestrians and motorcycles, then in FIG. 7A, the second object 720a, the third object 730a, and the fourth object 740a are removed, and only the first object 710a and the fifth object 750a remain.

[0087] FIG. 7B shows the variation data before filtering, similar to FIG. 7A, and FIG. 7B shows a diagram that a sixth object 710b has been recognized as an object.

[0088] More specifically, the sixth object 710b shown in FIG. 7B can be understood as not being recognized in the first recognition result but being recognized in the second recognition result, processed as variation data, and transmitted to the weak point analysis module 650. The weak point analysis module 650 can filter the variation data according to preset filter criteria to leave only meaningful object information in the variation data.

[0089] However, in Figure 7B, the sixth object 710b is not a single object, but rather a result of the seventh object 720b and the eighth object 730b accidentally overlapping and being mistakenly recognized as a single object, and due to its morphological characteristics, it has been recorded with a very low confidence of 0.3396.

[0090] 7B, the preset filter criterion is a reliability criterion regarding the reliability of the bounding boxes included in the variation data, and the weak point analysis module 650 may filter out bounding boxes with a reliability lower than the reliability criterion as information based on the variation data. Here, the reliability criterion may be 0.6, but may vary depending on the embodiment.

[0091] 7B, the weak point analysis module 650 can remove the bounding box of the sixth object 710b according to the confidence criterion. After the bounding box of the sixth object 710b is removed, there is no bounding box remaining in the frame of FIG. 7B, so the first recognition result and the second recognition result can be considered to be substantially the same. The fact that the first recognition result and the second recognition result are substantially the same means that the first model 620 does not need to learn the sixth object 710b.

[0092] Figure 7C shows the variability data before filtering, similar to Figures 7A and 7B, and Figure 7C shows diagrammatically that the ninth object 710c, the tenth object 720c, and the eleventh object 730c have been recognized as objects.

[0093] More specifically, of the objects shown in Figure 7C, the tenth object 720c and the eleventh object 730c are vehicles that were recognized as objects in both the first and second recognition results, and have had their bounding boxes removed. However, Figure 7C shows that the ninth object 710c is an object that is not likely to affect the driving of an autonomous vehicle traveling on a road, but has been classified as a truck and has had a bounding box applied to it.

[0094] Generally, the second model 630, which has higher recognition performance, recognizes a larger number of objects. However, in certain cases, the first model 620 may mistakenly recognize a non-object as an object, or the second model 630 may malfunction and mistakenly recognize an object not recognized by the first model 620 as a normal object. Therefore, the weak point analysis module 650 may determine that the ninth object 710c is an object that exists only on the road in a location that is not an actual road, and remove the corresponding bounding box according to preset filter criteria. In FIG. 7C, when the bounding box of the ninth object 710c is removed, the variation between the first and second recognition results is substantially eliminated, and naturally, the data for the first model 620 to learn is also eliminated.

[0095] FIG. 8 is a diagram illustrating a process of applying active learning to improve the object recognition rate of an autonomous vehicle according to another embodiment of the present invention.

[0096] The apparatus for improving an object recognition rate according to the present invention may include, in physical or logical forms, a classification module 820, a labeling data collection module 840, a learning model 850, and a prediction model 860, as shown in Fig. 8. In Fig. 8, the learning model 850 refers to a model that is being trained using input data, and the prediction model 860 refers to a predictive model that can output result data when test data is input after training is completed. Since the learning model 850 is a model whose recognition rate is improved through training, it ultimately refers to the first model 620 that is installed in an autonomous vehicle.

[0097] Typically, data labeling, which is an essential process in the process of preprocessing raw data for machine learning, is performed by humans because the features of the data cannot be accurately distinguished. However, the object recognition rate improvement device according to the present invention performs active labeling through active learning that partially includes auto-labeling, thereby enabling the learning model 850 to quickly and efficiently learn the features of the raw data 810.

[0098] In FIG. 8, raw data 810 refers to images captured and collected by a camera while the autonomous vehicle is traveling, similar to FIG.

[0099] The raw data 810 may be automatically labeled by the classification module 820. Specifically, if the raw data 810 is an image composed of multiple frames, the classification module 820 may automatically recognize objects for each frame and automatically classify the object classes, such as object a being a truck, object b being a pedestrian, and object c being a motorcycle, in a particular frame.

[0100] In analyzing raw data 810, classification module 820 does not automatically label objects determined to be difficult to classify by its internal classification algorithm. Objects determined to be difficult to classify may be the weak points described with reference to FIGS. 6 to 7C. That is, first object 710a and fifth object 750a in FIG. 7A, which are determined to be differences between the results of first model 620 and second model 630 even after filtering based on the filter criteria, may be objects determined to be difficult to classify by classification module 820. Information about objects determined to be difficult to classify is automatically collected by classification module 820 and transmitted to user 830, who has mastered advanced classification criteria. User 830 completes data labeling and then transmits labeling data 835 to labeling data collection module 840.

[0101] The labeling data collection module 840 receives all automatically labeled data from the classification module 820 and all manually labeled data from the user 830, and controls the learning model 850 to learn the labeled data. Data that cannot be learned in the learning model 850 due to irregularities is transmitted again to the classification module 820, where it is labeled by the classification module 820 or the user 830 and re-input into the learning model 850. This process is repeated until the model that has finally completed learning regarding object recognition of the raw data 810 becomes a prediction model 860, which can accurately recognize objects included in the newly input raw data 810.

[0102] As described above, by applying active learning in which only selected data is labeled by a user 830 who has mastered advanced classification standards and the remaining data is automatically labeled, the learning model 850 according to the present invention can quickly and accurately learn training data (information about objects in a video), and the classification module 820 applies the filter standards described in Figures 7A to 7C, thereby significantly reducing the amount of manual labeling work that the user 830 must do. In other words, according to the present invention, the excessive costs (time costs, monetary costs) incurred by existing labeling work can be minimized.

[0103] FIG. 9 is a flowchart showing an example of a method for improving an object recognition rate according to the present invention.

[0104] The method shown in FIG. 9 can be realized by the above-mentioned object recognition rate improving device, and therefore will be described below with reference to FIGS. 6 to 8, and description that overlaps with the content described with reference to FIGS. 6 to 8 will be omitted.

[0105] The object recognition rate improving device may recognize an object included in a first image acquired while driving using a first recognition technique and calculate a first recognition result (S910).

[0106] The object recognition rate improving apparatus may receive a second recognition result obtained by recognizing an object included in the first image using a second recognition technique (S930).

[0107] The object recognition rate improving device can calculate variation data of the first recognition result and the second recognition result (S950).

[0108] The object recognition rate improving device may control the first model for recognizing an object included in an image using a first recognition technique based on information on the variation data calculated in step S950 (S970).

[0109] FIG. 10 is a diagram illustrating a method for improving an object recognition rate according to still another embodiment of the present invention.

[0110] This alternative embodiment shares some of the same processes as the object recognition rate improvement method described with reference to Figures 6 to 9. The configuration for analyzing video captured during driving to recognize objects is similar, but unlike the method of Figure 6, which applies different recognition techniques to the same video to recognize objects and calculate variation data, this embodiment recognizes objects contained in the video using a single recognition technique. To distinguish it from the first model 620 and second model 630 described above, the model for recognizing objects in the video in this embodiment is referred to as a recognition model.

[0111] 10, a total of four frames are shown, and at least one object is located at a specific position in each frame. More specifically, in FIG. 10, objects are recognized to exist at the top and bottom of the i-th frame, the (i+1)-th frame, and the (i+3)-th frame, respectively, but in the (i+2)-th frame, the object at the bottom temporarily disappears, and it is recognized that only the object exists at the top. The object recognition rate improving device according to the present embodiment can regard a case where an object is recognized within a short time after the object suddenly disappears in a specific frame during tracking of a specific object as a weak point, as shown in FIG. 10, and convert it into learning data for training a recognition model.

[0112] In other words, this embodiment can be understood as an embodiment for improving object recognition performance by additional learning of the object recognition module, since the performance of the object recognition module of an autonomous vehicle is limited when an object that has been tracked normally disappears and then reappears in a specific frame.

[0113] [Table 4] Table 4 is a table listing the differences between the embodiment described using Figures 6 to 9 and the embodiment described in Figure 10. Referring to Table 4, it can be seen that the two embodiments of the present invention both have the same objective of identifying the point at which a performance limit (weakness point) occurs in an object recognition module installed in an autonomous vehicle, generating learning data to compensate for the identified performance limit, and quickly and efficiently training the object recognition module (recognition model), but there are some differences in the configurations used to achieve this.

[0114] FIG. 11 is a flowchart illustrating a method for improving an object recognition rate according to the embodiment described in FIG.

[0115] First, the object recognition rate improving device may recognize a first object from a first image captured while driving (S1110). Here, the object recognition rate improving device recognizing the first object from the first image means that the device recognizes the first object from a frame constituting the first image and obtains information about the size and class of the first object, as shown in FIG.

[0116] Next, the object recognition rate improving apparatus may detect whether the first object reappears after disappearing for a predetermined period of time in the first image (S1130).

[0117] Here, the predetermined period may be a time range value of at least one frame. If the frame rate of the collected first image is 30 frames per second, the predetermined period may be a time range value corresponding to 0 seconds to 1 / 30 seconds.

[0118] As another example, the predetermined period may be a time range value of 1 to 3 frames, and in Figure 10, it can be seen that the predetermined period is a time range value of 1 frame. If the predetermined period is a time range value of 3 frames, when a first object tracked in the i-th frame disappears in the i+1-th frame and reappears in the i+5-th frame, it can be considered to have disappeared for the predetermined period.

[0119] The object recognition rate improvement device can calculate learning data for the first object based on detecting that the first object has reappeared (S1150). If the first object does not reappear after disappearing, or if it reappears after a predetermined period of time has elapsed, the object recognition rate improvement device considers that the condition is not satisfied and does not calculate learning data for the first object. In particular, if the first object reappears after a time longer than the predetermined period of time has elapsed since disappearing, it is highly likely that the first object is not recognized due to the limitations of the recognition model's recognition performance, but rather is occluded by another object and therefore cannot be considered to satisfy the condition for calculating learning data.

[0120] In step S1150, the learning data may include at least one of the size, position, and class of the first object, information regarding the history of the first object disappearing for a predetermined period of time after being initially recognized and then reappearing, and information regarding the confidence of the first object.

[0121] The object recognition rate improvement device can control the learning of a recognition model for an autonomous vehicle that recognizes objects from images acquired while driving using information based on the learning data calculated in step S1150 (S1170).

[0122] In step S1170, the information based on the training data means information that has been further processed at least once so that the training data calculated in step S1150 can be input into a recognition model. For example, it may be information that has been filtered using preset filter criteria from the training data.

[0123] As an alternative embodiment, the preset filter criterion may be a filter criterion related to the length of time over a series of frames when a first object is recognized in a first frame, disappears in a second frame, and then reappears in a third frame, and the object recognition rate improving apparatus may calculate information based on learning data based on the filter criterion only if the length of time between the first frame and the third frame is longer than 10 frames. The filter criterion means that only objects that have been tracked for a sufficient length of time over several frames are selectively learned.

[0124] In another alternative embodiment, the preset filter criterion may be a classification criterion for classifying the type of the first object that is recognized in a first frame, disappears for a predetermined period in a second frame, and then reappears in a third frame, and the object recognition rate improving apparatus may calculate information based on learning data based on the classification criterion only when the type (class) of the first object is a car, a truck, a bus, or a miscellaneous object (misc.). The filter criterion means that learning is focused on cars, trucks, buses, and miscellaneous objects, which are objects that are highly important in autonomous driving.

[0125] In yet another alternative embodiment, the preset filter criterion may be a size criterion for classifying the size of the first object that is recognized in a first frame, disappears for a predetermined period in a second frame, and then reappears in a third frame, and the object recognition rate improving apparatus may calculate information based on learning data based on the size criterion if the height or width of the first object exceeds a preset number of pixels based on the size criterion. The filter criterion means that a recognition model is trained only for first objects that are sufficiently large.

[0126] As explained by the comparison in Table 4, when an object disappears and then reappears, the recognition model cannot recognize the object even though the object has not completely disappeared in the section where the object disappeared. This is due to the limited performance of the recognition model, so it can be classified as a weak point of the recognition model explained in Figure 8, and active learning can be applied in the same way.

[0127] That is, when the types of objects included in the training data are accurately labeled by input from a user who is familiar with the object classification criteria, the labeled data can be input to the recognition model as information based on the training data via a labeling data collection module. The recognition model that has completed learning through repeated learning can accurately recognize the second object in the second video without missing any frames when the second video is input as new test data.

[0128] FIG. 12 is a conceptual diagram conceptually illustrating an example of a method for generating improved training data images according to the present invention.

[0129] 12 conceptually illustrates the steps of the method for generating improved training data video according to the present invention, dividing them into sections for performing the steps, and intuitively illustrates the steps of acquiring video, recognizing objects from the video and detecting some frames, integrating the detected frames, and sampling the integrated frames. The process performed at each step will now be described in detail.

[0130] First, in step 1210, a camera mounted on the autonomous vehicle may collect images captured while the autonomous vehicle is traveling. The images collected in step 1210 may be video images having a certain frame rate and consisting of a number of frames. The images captured and generated by the camera mounted on the autonomous vehicle may be transmitted via wire or wirelessly to a device for generating improved learning data images according to the present invention. For convenience, the images captured while the autonomous vehicle is traveling may be referred to as first images.

[0131] In the present invention, training data video is video intended to train a recognition model that analyzes the video and recognizes objects in the video, and the higher the quality of the training data video, the more accurately the recognition model can recognize objects contained in the video, while the lower the quality of the training data video, the lower the recognition rate of objects contained in the video. In other words, the improved training data video generation method according to the present invention can provide a methodology for generating training data video of relatively higher quality compared to known training data.

[0132] Next, in step 1220, a Weakness Point Detection (WPD) Large Model (Large Model) may be applied to the collected first video to recognize objects contained in the first video and detect specific frames in which the objects are recognized. Here, the WPD Large Model may be a model included in a device that recognizes objects (such as buses, cars, trucks, pedestrians, and motorcycles) contained in a video using a unique recognition technique, distinguishes between objects that should be recognized but are not, and objects that should not be recognized but are recognized as objects, and processes the results to generate training data for improving the recognition rate of objects in the video. For example, the WPD Large Model of FIG. 12 may be a concept including the first model 620, the second model 630, the variation data calculation module 640, and the weak point analysis module 650 of FIG. 6.

[0133] In step 1230, WPD Tracking may be applied to the collected first video to recognize an object contained in the first video and detect a specific frame related to the object. Here, WPD Tracking may be a model that recognizes an object contained in the video using a unique recognition technique, similar to the WPD Large Model described above, and then analyzes the recognition results to detect a specific frame. For example, the WPD Tracking of FIG. 12 may be a model that implements the tracking algorithm described with reference to FIGS. 10 and 11. That is, when a specific object is recognized in the first video and tracking is initiated, if the object whose tracking has been initiated disappears temporarily (for several frames) and then reappears, the WPD Tracking of FIG. 12 may identify the frame numbers of frames where tracking has not been performed, calculate a result value, and add up the frames to generate training data video for improving the object recognition rate of the recognition model.

[0134] Hereinafter, a recognition technique is considered to refer to an algorithm that recognizes road objects (such as cars, trucks, buses, motorcycles, and pedestrians) that may affect the autonomous driving of an autonomous vehicle from a first video, and a detection technique is considered to refer to an algorithm that, once an object in the video is recognized by the recognition technique, detects a specific frame from among multiple frames that make up the video based on the results of analyzing the recognition result. For example, the first model 620 and the second model 630 in Figure 6 may be examples of recognition models that perform the recognition technique, and the WPD Large Model and WPD Tracking in Figure 12 may be examples of models that perform the detection technique.

[0135] Next, in step 1240, a process of collecting and integrating frames identified as a result of the WPD Large Model or WPD Tracking processing can be performed. Although learning data for improving the object recognition rate can be collected using the detection techniques performed in steps 1220 and 1230, the WPD Large Model-based detection technique has a limitation in that parts that are not recognized as objects by either of the two recognition models used are not continuously recognized, and the WPD Tracking-based detection technique has a limitation in that parts that are not initially recognized as objects are excluded from tracking from the beginning. Therefore, the present invention proposes a method of overcoming these technical limitations by integrating the result data of different detection techniques.

[0136] For example, if the frame numbers of the first image detected by the WPD Large Model are 1, 5, 14, 16, 32, and 50, and the frame numbers of the first image detected by WPD Tracking are 14 and 52, in step 1240, the numbers of the integrated frames will be 1, 5, 14, 16, 32, 50, and 52, and hereinafter the integrated frames will be collectively referred to as a frame set.

[0137] A sampling process can be performed on the integrated frames in step 1250. The sampling in step 1250 takes into consideration two points: since frames are selected from the same video by changing only the detection technique, very similar information is contained between consecutive frames; and if too many frames are detected in steps 1220 and 1230, the acquired "learning data video" contains unnecessarily large amounts of information, which may result in overfitting of the object recognition rate.

[0138] In the present invention, various methods for sampling the multiple integrated frames are available. For example, a method may be used in which frames included in each time interval are randomly sampled at predetermined time intervals. When the integrated frame numbers are 1, 5, 14, 16, 32, and 50 at a constant frame rate, sampling every 10 frames may result in a total of four frames being sampled, with either one of frames 1 and 5, and either one of frames 14 and 16 sampled along with frames 32 and 50. The device according to the present invention may generate an improved training data video by combining the four sampled frames in chronological order. The training data video generated by the device according to the present invention through a series of processes on the first video is also referred to as the second video, and the device according to the present invention will hereinafter be referred to as the second video generation device. The second video generation device may also be implemented in a form that physically or logically includes the object recognition rate improvement device described with reference to FIGS. 5 to 11.

[0139] In one embodiment, a frame group may be used to sample frames included in a frame set in step 1250. More specifically, the second image generating device may generate frame groups including at least one frame and including non-overlapping frames based on the frame set generated in step 1240, then extract frames for each frame group, and combine the extracted frames to generate the second image.

[0140] [Table 5] Table 5 illustrates an embodiment in which frame groups are generated from a frame set to generate a second image. In Table 5, the frame numbers detected by the first detection technique are 1, 5, 14, 16, 32, and 51. Here, the first detection technique may be, but is not limited to, a detection technique based on the WPD Large Model of FIG. 12. Also, in Table 5, the frame numbers detected by the second detection technique are 14 and 53. Here, the second detection technique may be, but is not limited to, a detection technique based on the WPD Tracking of FIG. 12. As described above, the first and second detection techniques may detect frames using different algorithms. When the detected frames are combined in step 1240 of FIG. 12, the frame numbers included in the frame set in Table 5 are found to be 1, 5, 14, 16, 32, 51, and 53. In this embodiment, the second image generation device samples the frames included in the frame set into groups of 10 frames to generate a total of four frame groups. In this embodiment, when four frame groups are generated as shown in Table 5, the second image generating device can extract frames for each frame group to generate the second image.

[0141] Depending on the embodiment, the second image generating device may extract one frame for each frame group to generate the second image, or may extract several frames for each frame group based on a weight set for each frame group or a weight set for each frame belonging to a frame group to generate the second image.

[0142] For example, the weight set for each frame group may be a value determined based on the number of frames included in each frame group. If frame group A includes 10 frames and frame group B includes 5 frames, the second image generation device may perform sampling by selecting two frames from frame group A and one frame from frame group B based on the weight ratio of frame groups A and B being 2:1.

[0143] As another example, the second image generating device may extract frames from each frame group by considering the weight assigned to a specific frame belonging to the frame group. In Table 5, the frame number detected by both the first and second detection techniques is 14, and frame 14 of the first image has a higher weight than the other frames. That is, weights can be assigned as metadata for each frame, and frames that are repeatedly detected by various detection techniques can be assigned a higher weight. In particular, while FIG. 12 and Table 5 illustrate two detection techniques, if the number of detection techniques increases to three or more as shown in FIG. 13 (described below), frames included in a frame set can be assigned weights of various magnitudes. The second image generating device samples frames from several frame groups to necessarily select frames with high weights, resulting in the second image including frames with high weights.

[0144] In yet another embodiment, the weight assigned to a specific frame belonging to a frame group may be a value dependent on the weight assigned to each detection technique. For example, assuming that a first, second, and third detection technique are present, and the weight of the first detection technique is 1, the weight of the second detection technique is 2, and the weight of the third detection technique is 3, the weight of a frame detected by the first and second detection techniques is lower than the weight of a frame detected by the first and third detection techniques, which in turn is lower than the weight of a frame detected by the second and third detection techniques. In other words, in the present invention, frames included in a frame set have, as metadata, information on the weight ultimately determined by not only the frame group to which each frame belongs but also the value assigned to each detection technique by which each frame was detected. Such weights can be an effective factor in the sampling process performed by the second image generation device.

[0145] FIG. 13 is a diagrammatic illustration of an expanded concept of the embodiment described in FIG.

[0146] Comparing Figure 13 with Figure 12, it can be seen that they have in common the fact that a first image acquired during driving is input and becomes input data for a second image generating device (step 1310), and that frames detected using multiple detection techniques are integrated into a frame set (step 1340) and then sampled after integration (step 1350). The only difference is the process (1320A-1330B) in which an object recognition algorithm is applied to the input first image and frames in which an object is recognized are detected.

[0147] In particular, in Figure 12, there are two entities (WPD Large Model and WPD Tracking) that perform detection techniques to detect the frames that form the basis of the second image, but Figure 13 diagrammatically shows that there can be four or more entities.

[0148] 13, Large Model 1 and Large Model 2 refer to recognition models related to the WPD Large Model that process a recognition technique for recognizing an object included in a first image using a unique recognition technique, as described above, and are characterized in that the combination of recognition techniques used varies depending on the identification number. For example, Large Model 1 may be a detection technique that detects frames using an algorithm for recognizing an object in a first image based on YoloV4-CSP described in Table 2 and an algorithm for recognizing an object in a first image based on YoloV4-P7 described in Table 3, and Large Model 2 may be a detection technique that detects frames using an algorithm for recognizing an object in a first image based on YoloV4-P5 and an algorithm for recognizing an object in a first image based on YoloV4-P6 described in Table 3.

[0149] Also, in Figure 13, Tracking 1 and Tracking 2 refer to different versions of the WPD tracking algorithm described in Figures 10 and 11. In particular, in the case of WPD Tracking, various versions may be available depending on the time criteria for when a once-recognized object disappears for a predetermined period of time and then reappears. For example, in Figure 13, Tracking 1 may be a detection technique that detects frames by regarding a once-recognized object that has disappeared for three frames and then reappears as a weak point, and Tracking 2 may be a detection technique that detects frames by regarding a once-recognized object that has disappeared for two frames and then reappears.

[0150] Although only the WPD Large Model and tracking algorithm are described in FIG. 13, various object recognition algorithms such as SORT (Simple Online and Realtime Tracking), Bytetrack, and StrongSORT can be further added depending on the embodiment.

[0151] That is, in Fig. 13, the number of techniques for recognizing an object in the first image and detecting frames of the recognized object in various ways is not limited depending on the basic detection algorithm, learning data, etc. used by each detection technique, so although a total of four detection techniques are shown in Fig. 13, fewer or more detection techniques than four may be used depending on the embodiment. Also, similar to the embodiment described in Fig. 12, even if the number of detection techniques is four or more, all of the various embodiments described in Fig. 12 may be applied.

[0152] FIG. 14 is a flowchart showing an example of the second video generating method according to the present invention.

[0153] The method according to Figure 14 can be realized by the second image generation device described in Figures 12 and 13, so the following description will be given with reference to Figures 12 and 13, and any overlapping description with that already described will be omitted.

[0154] The second image generating device may apply at least two or more recognition techniques to a first image acquired while driving to recognize an object included in the first image (S1410).

[0155] The second image generation device may apply at least two or more detection techniques to the object recognition result in step S1410 and detect frames for each detection technique (S1430).

[0156] The second image generating device may generate a frame set including a plurality of frames by integrating the frames detected in step S1430 (S1450).

[0157] The second image generating device may generate a second image by sampling the merged frame set (S1470).

[0158] The second image generated in step S1470 can be high-quality training data that is more useful for improving object recognition rates compared to training data generated by existing methods by diversifying detection techniques and sampling according to unique features.

[0159] FIG. 15 is a block diagram of a second image generating device according to an embodiment.

[0160] 15, a second image generation device 1500 may include a communication unit 1510, a processor 1520, and a DB 1530. Only components related to the embodiment are shown in the second image generation device 1500 in Fig. 15. Therefore, a person skilled in the art would understand that the second image generation device 1500 may further include other general-purpose components in addition to the components shown in Fig. 15.

[0161] The communication unit 1510 may include one or more components that enable wired / wireless communication with an external server or device. For example, the communication unit 1510 may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast receiving unit (not shown).

[0162] The DB 1530 is hardware that stores various data processed within the second image generating device 1500, and can store programs for processing and controlling the processor 1520.

[0163] DB1530 may include random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disc storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.

[0164] The processor 1520 controls the overall operation of the second image generation device 1500. For example, the processor 1520 may execute a program stored in the DB 1530 to overall control an input unit (not shown), a display (not shown), a communication unit 1510, the DB 1530, etc. The processor 1520 may control the operation of the second image generation device 1500 by executing a program stored in the DB 1530.

[0165] The processor 1520 can control at least a part of the operation of the second image generation device 1500 described above with reference to FIGS.

[0166] As an example, as described in Figures 6 to 9, the processor 1520 may recognize an object included in a first image acquired while the vehicle is traveling using a first recognition technique to calculate a first recognition result, receive a second recognition result obtained by recognizing an object included in the first image using a second recognition technique, calculate variation data of the first recognition result and the second recognition result, and control the first model operating using the first recognition technique to be learned using information based on the calculated variation data.

[0167] As another example, as described in Figures 10 and 11, processor 1520 can recognize a first object from a first image captured while driving, detect that the first object disappears for a predetermined period of time and then reappears in the first image, calculate learning data regarding the first object when it detects that the first object has reappeared, and control the system so that a recognition model that recognizes the object included in the image is learned using information based on the calculated learning data.

[0168] As yet another example, the processor 1520 may apply at least two or more recognition techniques to a first image acquired while driving to recognize objects contained in the first image, apply at least two or more detection techniques to the results of the object recognition to detect frames for each detection technique, integrate the detected frames to generate a frame set including multiple frames, and sample the integrated frame set to generate a second image.

[0169] The processor 1520 may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.

[0170] FIG. 16 is a flowchart illustrating a method for improving an object recognition rate according to the embodiment described in FIG.

[0171] First, the object recognition rate improving device may recognize a first object from a first image acquired while driving (S1610). Here, the object recognition rate improving device recognizing the first object from the first image means that the device recognizes the first object from a frame constituting the first image and obtains information about the size and class of the first object, as shown in FIG.

[0172] Next, the object recognition rate improving apparatus may detect whether the first object disappears for a predetermined period of time and then reappears in the first image (first frame) (S1630).

[0173] Here, the predetermined period may be a time range value of at least one frame. If the frame rate of the collected first image is 30 frames per second, the predetermined period may be a time range value corresponding to 0 seconds to 1 / 30 seconds.

[0174] As another example, the predetermined period may be a time range value of 1 to 3 frames, and in Figure 16, it can be seen that the predetermined period is a time range value of 1 frame. If the predetermined period is a time range value of 3 frames, when a first object tracked in the i-th frame disappears in the i+1-th frame and then reappears in the i+5-th frame, it can be considered to have disappeared for the predetermined period.

[0175] The object recognition rate improving device can calculate learning data related to the first object based on the detection of the reappearance of the first object (S1650). Furthermore, based on the detection of the reappearance of the first object, the object recognition rate improving device can consider the first object to have been recognized in the previous frame in which it was determined that the first object had disappeared for a predetermined period and was not recognized.

[0176] For example, in FIG. 6, if a first object is recognized in the i-th frame and the i+1-th frame, and then disappears in the i+2-th frame and is recognized again in the i+3-th frame, the object recognition device (object recognition module) included in the object recognition rate improvement device still cannot recognize the first object in the i+2-th frame, but it is considered that the first object recognized in the i-th frame, the i+1-th frame, and the i+3-th frame are also present in the i+2-th frame.

[0177] If the first object does not reappear after disappearing, or if it reappears after a predetermined period of time has passed, the object recognition rate improvement device considers that the condition is not satisfied and does not calculate learning data for the first object. In particular, if the first object reappears after a time longer than the predetermined period of time has passed after disappearing, it is not because the recognition model is unable to recognize the first object due to limitations in its recognition performance, but because the first object is likely to be occluded by another object and therefore unable to be recognized, and therefore the condition for calculating learning data cannot be considered to be satisfied.

[0178] In step S1650, the learning data may include at least one of the size, position, and class of the first object, information regarding the history of the first object disappearing for a predetermined period of time after being initially recognized and then reappearing, and information regarding the confidence of the first object.

[0179] The object recognition rate improvement device can control the learning of a recognition model for an autonomous vehicle that recognizes objects from images acquired while driving using information based on the learning data calculated in step S1650 (S1670).

[0180] In step S1670, the information based on the training data means information that has been further processed at least once so that the training data calculated in step S1650 can be input into a recognition model. For example, it may be information that has been filtered using preset filter criteria from the training data.

[0181] The tracking algorithm described in Fig. 16 may be implemented as a Kalman filter-based SORT (Simple Online and Realtime Tracking) algorithm, but is not limited to this. In particular, a tracking algorithm based on a Kalman filter has the property of operating correctly on a 2D bounding box and providing information on the results of the tracking algorithm in the form of a 2D bounding box, and the tracking algorithm described in the present invention can have the same property.

[0182] As an alternative embodiment, the preset filter criterion may be a filter criterion related to the length of time over a series of frames when a first object is recognized in a first frame, disappears in a second frame, and then reappears in a third frame, and the object recognition rate improving apparatus may calculate information based on learning data based on the filter criterion only if the length of time between the first frame and the third frame is longer than 10 frames. The filter criterion means that only objects that have been tracked for a sufficient length of time over several frames are selectively learned.

[0183] In another alternative embodiment, the preset filter criterion may be a size criterion for classifying the size of the first object that is recognized in a first frame, disappears for a predetermined period in a second frame, and then reappears in a third frame, and the object recognition apparatus may calculate information based on learning data based on the size criterion if the height or width of the first object exceeds a preset number of pixels based on the size criterion. The filter criterion means that a recognition model is trained only for first objects of a sufficiently large size.

[0184] FIG. 17 is a diagram for explaining the cuboid of an object in an image according to the present invention.

[0185] 17, the object recognition rate improving apparatus recognizes two objects from an image, and then generates outlines 1710a and 1730a for each object as a result of the recognition. The object recognition rate improving apparatus can generate a cuboid for each object based on the generated outlines 1710a and 1730a. That is, the outline 1710a of the motorcycle (two-wheeler) can be basic information for generating the motorcycle cuboid 1710b, and the outline 1730a of the car can be basic information for generating the car cuboid 1730b.

[0186] In the present invention, a cuboid may be generated in the form of two polygons joined together based on a common edge, one of which may represent the front or rear of an object, and the other may represent a side of the object. Depending on the embodiment, the cuboid may be displayed as a rectangle on the front and rear of an object, and a trapezoid on the left and right sides with parallel sides of different lengths, as shown in Fig. 17.

[0187] [Table 6] Table 6 shows an example of a method for analyzing the movement of an object when the cuboid is formed by two polygons joined horizontally along a common edge. As shown in Table 6, the cuboid of an object can be understood as object metadata that can intuitively and efficiently indicate the object's overall size, movement direction, and relative position with respect to the position of the camera capturing the image. As an example, the coordinate values ​​of the cuboid may be composed of seven coordinate values. In FIG. 17, the cuboid 1710b of the motorcycle has a shape in which a rectangle representing the front of the motorcycle and a trapezoid representing the side of the motorcycle are joined together along a common edge. Here, the vertex coordinates for forming the bike cuboid 1710b are (x1, y1) which is the coordinate of the top left edge, (x2, y1) which is the coordinate of the top center edge, (x3, y3) which is the coordinate of the top right edge, (x1, y2) which is the coordinate of the bottom left edge, (x2, y2) which is the coordinate of the bottom center edge, and (x3, y4) which is the coordinate of the bottom right edge, and the coordinate values ​​that are the minimum information required to form the bike cuboid 1710b are a total of seven (x1, x2, x3, y1, y2, y3, y4).

[0188] As another example, a cuboid may be composed of eight coordinate values. Although not shown in Figure 17, if an object is surrounded by a second outline of a rectangular parallelepiped shape to enhance the perspective of the object, a total of eight coordinates are required to form the cuboid of the object, and the total required coordinate values ​​are eight (x1, x2, x3, x4, y1, y2, y3, y4).

[0189] As shown in Figure 17, an outline in the form of a 2D bounding box may be generated for an object immediately recognized from an image (or frame), and a cuboid may be generated based on the coordinates of that outline, but if an object is not recognized in a frame by the tracking algorithm described in Figures 10 and 16, but is deemed to have been recognized in that frame as a result of applying the tracking algorithm, it still does not mean that the outline of the object has been generated in that frame. In other words, it is necessary to obtain a cuboid even for an object in a frame in which an object is not recognized due to limitations in the object recognition performance of the object recognition rate improvement device (object recognition device), but in which the object should have been recognized otherwise; a specific method for this will be described later using Figures 18 and 19.

[0190] FIG. 18 is a schematic diagram for explaining the method for obtaining the cuboid of an object in an image according to the present invention.

[0191] Hereinafter, an apparatus for realizing the method for acquiring cuboids of an object in an image according to the present invention will be referred to as a "cuboid acquisition apparatus" for short.

[0192] First, the cuboid acquisition device receives an input of an image (S1810). The image input (received) to the cuboid acquisition device in step S1810 includes at least two frames and may be an image captured by a camera mounted on a vehicle while the vehicle is traveling.

[0193] The cuboid acquisition device can recognize objects by applying an object recognition process to the video received in step S1810 (S1820). Here, as already explained, the object recognition process may be performed by an object detector.

[0194] The cuboid acquisition device may apply a tracking algorithm to the image received in step S1810 to determine whether an object should be recognized in a frame in which no object is recognized (S1830). Here, the tracking algorithm may be the Kalman filter-based SORT algorithm described above, or an algorithm other than SORT may be used depending on the embodiment.

[0195] The cuboid acquisition device can associate the results of steps S1820 and S1830 (S1840). The reason for associating the results of the object recognition process and the tracking algorithm in step S1840 is to determine when an object has not been detected (recognized) in the object recognition process but has been detected by the tracking algorithm. In the association process in step S1840, frames that are not necessary for determining the presence or absence of an object in step S1850, which will be described later, can be excluded.

[0196] The cuboid acquisition device can determine the presence or absence of objects that should have been recognized immediately by the object recognition process but were not recognized and were recognized only by the tracking algorithm (S1850) while integrating the results of steps S1820 and S1830. Objects determined in step S1850 may be classified as missed objects and assigned additional metadata, and a cuboid conversion process is performed in step S1860, which will be described later. As an example, a Hungarian algorithm can be applied in step S1850.

[0197] If it is determined in step S1850 that there is a missing object, the cuboid acquisition device can generate the cuboid coordinates of the object (S1860). If it is determined in step S1850 that there is no missing object, the function of the cuboid acquisition device can end without performing another cuboid transformation (S1870).

[0198] [Table 7] Table 7 is a table showing a concept for more specifically explaining the content described with reference to Figure 18. For convenience, the first frame and the second frame will be considered to be one of multiple frames included in one video, and the second frame will be considered to be the frame located immediately after the first frame. First, Table 7 shows the results where an object is immediately recognized by the object detector in the first frame, but the object is not recognized by the object detector in the second frame.

[0199] Next, Table 7 shows the results when an object is recognized immediately in the first frame and is considered to be recognized by the tracking algorithm in the second frame. Comparing the two results above, we can see that the object should have been recognized in the second frame of Table 7, but was not, and was recognized only by the tracking algorithm. Although not shown in Table 7, there is a third frame located after the second frame, and since the object is recognized in the third frame, the tracking algorithm is applied, and the object is considered to be recognized in the second frame.

[0200] In Table 7, in the first frame in which an object is immediately recognized, the outline and cuboid surrounding the object may be generated by the cuboid acquisition device according to the present invention. Once the outline (2D bounding box) surrounding the object is generated, the cuboid acquisition device can generate a cuboid for each object using the coordinate values ​​of the 2D bounding box, as described in Table 6.

[0201] On the other hand, in Table 7, in the second frame where the object is not immediately recognized and is considered to be recognized, the outline is not immediately generated, and the cuboid acquisition device indirectly acquires the coordinates of the outline of the object considered to be recognized in the second frame by using the coordinate values ​​of the outline generated in the previous and next frames (such as the first and third frames) based on the second frame.

[0202] Finally, as shown in Table 7, the cuboid of an object deemed to be recognized in the second frame cannot be obtained because its outline is not generated in the second frame. In this invention, a method is proposed for obtaining the coordinates of the cuboid of an object deemed to be recognized in the second frame by comprehensively considering the outline of the object deemed to be recognized in the first frame, the outline of the object deemed to be recognized in the second frame, and the coordinates of the cuboid of the object deemed to be recognized in the first frame, which will be described later using Figures 19 and 20.

[0203] FIG. 19 is a diagram for explaining the linear mapping method performed by the cuboid acquisition device according to the present invention.

[0204] For the sake of convenience, the following description will be made with reference to Table 7. The outline and cuboid in the first frame will be referred to as the first outline and first cuboid, and the outline and cuboid in the second frame will be referred to as the second outline and second cuboid.

[0205] The cuboid acquisition device according to the present invention can calculate a transform value using the coordinate values ​​of the first and second contour lines that have already been determined. As an example, the transform value may be an affine transform matrix. Referring to FIG. 19, when an affine transform is applied to the coordinates of the three points on the left side of FIG. 19, it can be seen that they are transformed into the coordinates of the three points on the right side of FIG. 19. The affine transform matrix can be expressed as a matrix related to translation, scaling, shear, and rotation.

[0206]

number

[0207]

number

[0208]

number

[0209]

number

[0210]

number

[0211] The cuboid acquisition device according to the present invention can acquire the second cuboid in the second frame using formulas 1 to 5.

[0212] FIG. 20 is a diagram for explaining another example of the linear mapping method performed by the cuboid acquisition device according to the present invention.

[0213] Referring to Figure 20, it can be seen that when a perspective transform is applied to the coordinates of the four points on the left side of Figure 20, they are converted into the coordinates of the four points on the right side of Figure 20. Referring to Figures 19 and 20, it can be seen that when at least three coordinate values ​​of the first outline and the second outline are given and the coordinate values ​​of the first cuboid are given, the coordinate values ​​of the second cuboid can also be obtained by the cuboid obtaining device according to the present invention.

[0214] FIG. 21 is a flowchart showing an example of a cuboid obtaining method according to the present invention.

[0215] The method according to Figure 21 can be realized by the cuboid acquisition device described with reference to Figures 17 to 20, and therefore will be described below with reference to Figures 17 to 20, and explanations that overlap with those already described will be omitted.

[0216] The cuboid acquisition device can recognize an object from the video and identify a first frame in which the object is recognized and a second frame in which the object is not recognized (S2110).

[0217] The cuboid acquisition device can generate a first outline of the object in the first frame and acquire coordinates of the first cuboid of the object based on first coordinate values ​​that form the first outline (S2130).

[0218] The cuboid acquisition device may determine whether an object should be recognized in the second frame and generate a second outline of the object in the second frame based on the determination result (S2150). As described above with reference to Table 3, in step S2150, the cuboid acquisition device may identify an object (or the outline of the object) recognized in frames surrounding the second frame. Here, the surrounding frames may be frames located before or after the second frame, such as the first and third frames. In addition, the number of surrounding frames and the frame numbers may vary depending on how long a predetermined period is set as for a tracking algorithm applied in the present invention when a recognized object disappears for a predetermined period and then reappears.

[0219] The cuboid obtaining device may calculate a conversion relation value between the first coordinate values ​​constituting the first outline and the second coordinate values ​​constituting the second outline (S2170).

[0220] The cuboid acquisition unit can calculate the coordinates of the second cuboid of the object in the second frame by applying the transformation relation value to the coordinates of the first cuboid (S2190).

[0221] FIG. 22 is a block diagram of a cuboid acquisition device according to one embodiment.

[0222] 22, a cuboid acquisition device 2200 may include a communication unit 2210, a processor 2220, and a DB 2230. Only components related to the embodiment are shown in the cuboid acquisition device 2200 in Fig. 22. Therefore, a person skilled in the art would understand that the cuboid acquisition device 2200 may further include other general-purpose components in addition to the components shown in Fig. 22.

[0223] The communication unit 2210 may include one or more components that enable wired / wireless communication with an external server or device. For example, the communication unit 2210 may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast receiving unit (not shown).

[0224] The DB 2230 is hardware that stores various data processed within the cuboid acquisition device 2200, and can store programs for processing and controlling the processor 2220.

[0225] DB2230 may include random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disc storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.

[0226] The processor 2220 controls the overall operation of the cuboid acquisition device 2200. For example, the processor 2220 can generally control the input unit (not shown), the display (not shown), the communication unit 2210, the DB 2230, etc. by executing a program stored in the DB 2230. The processor 2220 can control the operation of the cuboid acquisition device 2200 by executing a program stored in the DB 2230.

[0227] The processor 2220 may control at least a portion of the operation of the cuboid acquisition device 2200 described above.

[0228] As an example, as described in Figures 17 to 21, the processor 2220 can recognize an object from an image, identify a first frame in which the object is recognized and a second frame in which the object is not recognized, generate a first outline of the object in the first frame, obtain coordinates of a first cuboid of the object based on first coordinate values ​​constituting the first outline, determine whether the object should be recognized in the second frame, generate a second outline of the object in the second frame based on the determination result, calculate a conversion relationship value between the first coordinate values ​​constituting the first outline and the second coordinate values ​​constituting the second outline, and then apply the conversion relationship value to the coordinates of the first cuboid to calculate the coordinates of the second cuboid of the object in the second frame.

[0229] The processor 2220 may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.

[0230] FIG. 23 is a schematic diagram for explaining the method for improving the object recognition rate according to the present invention.

[0231] Hereinafter, a device that implements the method according to the present invention will be referred to as an "object recognition rate improving device."

[0232] First, the apparatus for improving object recognition rate receives an input of a video (S2310). The video input (received) to the apparatus in step S2310 may be a video including at least two frames captured by a camera mounted on a vehicle while the vehicle is traveling.

[0233] The object recognition rate improving device may recognize objects by applying an object recognition process to the video received in step S2310 (S2320). Here, the object recognition process may be performed by an object detector included in the object recognition rate improving device. The object recognition device may selectively recognize, as objects, only the main objects (e.g., passenger cars, trucks, buses, motorcycles, humans, and unidentified objects) necessary for realizing the autonomous driving function of a vehicle from among various objects included in a video frame using an object recognition algorithm configured therein. As described with reference to FIGS. 4 and 5, the object recognition algorithm of the object recognition device is intuitive and convenient in that it can immediately recognize some necessary objects from various objects in a video. However, due to performance limitations of the object recognition algorithm, it may not be possible to recognize an object that should be recognized as an object. If the object recognition algorithm is unable to recognize an object that should be recognized as an object due to performance limitations, the tracking algorithm described with reference to FIGS. 10 and 16 may be used to compensate to a certain extent.

[0234] The object recognition rate improving device may apply the tracking algorithm described with reference to Figures 10 and 16 to the video received in step S2310 to determine whether an object should be recognized in a frame in which no object is recognized (S2330). Here, the tracking algorithm may be the Kalman filter-based SORT algorithm described above, or an algorithm other than SORT may be used depending on the embodiment.

[0235] The object recognition rate improving apparatus may process the tracking results into history information while performing the tracking algorithm in step S2330 and store the history information in a history database (S2340). The history database may store not only the tracking results but also the results processed in steps S2350, S2370, and S2380, which will be described later.

[0236] The object recognition rate improving apparatus may associate the results of steps S2320 and S2330 (S2350). The reason for associating the results of the object recognition process and the tracking algorithm in step S2350 is to determine if an object is not detected (recognized) in the object recognition process but is detected only by the tracking algorithm. In step S2350, the object recognition result in step S2320 and the tracking result in step S2330 may be matched using a Hungarian algorithm. Specifically, in step S2350, it is determined whether the object outline (2D bounding box) generated for each frame as the object recognition result in step S2320 matches the object outline (2D bounding box) generated for each frame as the tracking result in step S2330. The matched outlines for each frame may have a match flag set to true and may be stored in a history database as history information, as shown in FIG. 23. In the integration process in step S2350, frames that are not necessary for determining the presence or absence of an object (frames in which the outline of a recognized object is not detected) can be excluded.

[0237] If the object recognition rate improving device cannot detect an object in chronologically consecutive frames, it deletes the track of the object (stops tracking) and may refer to the history information stored in the history database (S2360). In step S2360, the object recognition rate improving device may check the history information stored in the history database in reverse order, check the match flag of the outline of the object, and delete (exclude) frames in which the match flag is false consecutively from all tracks. Step S2360 will be described with reference to FIG. 25.

[0238] Next, the object recognition rate improving device can check whether there are any unmatched outlines based on the results of the matching operation in step S2350 (S2370). Checking whether there are any unmatched outlines in step S2370 means determining whether there are any outlines of objects that have been detected only by the tracking algorithm in step S2330. If there are any such outlines, the object recognition rate improving device can set the match flag to false, as in step S2360, and store the history information in the history database.

[0239] The object recognition rate improving device can perform post-track registration processing based on the results of the determinations made in steps S2350 to S2370 (S2380). The object recognition rate improving device performs data processing by referring to the information stored in the history database in step S2380, and specifically, can utilize information on whether the match flag for the outline of each object stored in the history database is true or false. In addition, the object recognition rate improving device can perform post-track registration processing by referring to information on the first and second reference values, which are the setting values ​​of the first and second parameters. Step S2380 will be described later with reference to FIG. 24.

[0240] FIG. 24 is a diagram for explaining the processing process after track registration explained in step S2380 of FIG.

[0241] First, referring to the upper part of Figure 24, it can be seen that the object recognition rate improving device recognizes objects by applying an object recognition algorithm to eight frames arranged in chronological order, resulting in object recognition in a total of five frames. The frames in which objects are recognized are frames t3, t4, t6, t7, and t8. Although objects are also present in the other three frames (frames t1, t2, and t5), it is assumed that the objects were not recognized by the object recognition algorithm of the object recognition rate improving device due to issues such as lighting, angle, distance, or color, or that the objects were not recognized due to a malfunction of the object recognition algorithm. That is, the upper part of Figure 24 shows the result of step S2320 in Figure 23, and an outline in the form of a 2D bounding box is generated for the object in the frames in which the object is recognized.

[0242] Next, referring to the bottom of Figure 24, it can be seen that an object is recognized in frame t8 as a result of the tracking algorithm performed by the object recognition rate improving device. The first reference value set in the object recognition rate improving device is 3, and the object is detected consecutively in frames t6, t7, and t8, satisfying the first reference value, and thus the track of the object is registered in frame t8. Because the track of the object is registered in frame t8 according to the tracking algorithm described with reference to Figures 10 and 16, even if the object disappears in frame t9 (not shown in Figure 24), tracking of the object can be maintained for the number of frames determined by the second reference value.

[0243] On the other hand, since the first reference value is set to 3 in the object recognition rate improving device of Figure 24, the track of the object is registered at frame t8 in Figure 24, and the start of the track is frame t6. However, since the object recognition rate improving device according to the present invention includes a process for minimizing problems that occur when the first parameter and the second parameter are changed to different values, it is possible to use the match flag (true or false) stored in the history database to check whether there is an object that has not been recognized even in frames before the track is registered.

[0244] Referring to FIG. 24, the object recognition rate improving device according to the present invention refers to information stored in a history database, checks the object recognition results and tracking results of frames prior to frame t6, where tracking begins, and can determine that an object is recognized in frames t4 and t6, but is temporarily not recognized in frame t5.

[0245] Since the first reference value of the object recognition rate improving device is 3, a track is not registered simply because an object is recognized twice consecutively in frames t3 and t4. However, after a track is registered, the object recognition rate improving device according to the present invention can check history information in previous frames and, if there is an object that should have been recognized but has not been recognized because the first reference value has increased to 3, can operate to include the object in the recognition result. That is, in Figure 24, the object recognition rate improving device can include all of the objects in frames t3, t4, t6, and t7 in the tracking result, determine that the object has been recognized by the tracking algorithm, and process so that history information corresponding to the outline matching result (false match flag) in frame t5 is stored in the history database. The above process can be processed in step S2380 described above.

[0246] FIG. 25 is a diagram for explaining the processing process after the track deletion described in step S2360 of FIG.

[0247] First, referring to the upper part of Figure 25, it can be seen that the object recognition rate improving device recognizes an object in a total of one frame by applying an object recognition algorithm to four frames arranged in chronological order to recognize the object. The frame in which the object is recognized is frame t1, and the object is not recognized in the other three frames (frames t2, t3, and t4). In the upper part of Figure 25, a first outline 2510A in the form of a 2D bounding box is generated for the object recognized by the object recognition algorithm.

[0248] Next, referring to the bottom of Figure 25, it can be seen that an object is detected consecutively in frames t1, t2, and t3 as a result of the tracking algorithm performed by the object recognition rate improving device. The track of the object has already been registered, and at t1, the object recognition result and the tracking result show the same outline, and the match flag is true. Therefore, even if the object disappears in frame t2, tracking of the object continues for the number of frames determined by the second reference value, and outlines 2530 and 2550 corresponding to the object can also be detected, as shown in the bottom of Figure 25.

[0249] Since the second reference value is set to 2 in the object recognition rate improving apparatus of Fig. 25, the outline of the object is still detected in frames t2 and t3 in Fig. 25. However, since the object recognition rate improving apparatus according to the present invention includes a process for minimizing problems that occur when the first parameter and the second parameter are changed to different values, it is possible to exclude frames t2 and t3 from all tracks by utilizing the match flag (false) stored in the history database. The results of the object recognition rate improving apparatus according to the present invention can also be stored as history information in the history database.

[0250] FIG. 26 is a flowchart showing an example of a method for improving an object recognition rate according to the present invention.

[0251] The method shown in Figure 26 can be realized by the object recognition rate improvement device described with reference to Figures 10, 16, 23, 24, and 25, and therefore will be described below with reference to Figures 10, 16, 23, 24, and 25, and any explanation that overlaps with what has already been described will be omitted.

[0252] The object recognition rate improving apparatus may recognize an object included in an image using a first recognition algorithm that recognizes an object for each frame of the image (S2610). In step S2610, the first recognition algorithm refers to an object recognition algorithm by an object detector, and may correspond to step S2320 in FIG. 23.

[0253] The object recognition rate improving device may form a track using a plurality of frames included in the video and recognize the object using a second recognition algorithm that recognizes the object included in the track (S2630). In step S2630, the second recognition algorithm refers to a tracking algorithm used by the object recognition rate improving device, and may correspond to step S2330 in FIG. 23.

[0254] In an alternative embodiment, the second recognition algorithm may be an algorithm that selectively recognizes an object that has been recognized in consecutive frames equal to or greater than a first threshold, then disappears for a number of frames equal to or less than a second threshold, and then reappears, as described above with reference to the first and second parameters.

[0255] The apparatus for improving object recognition rates may compare the result of recognizing an object using the first recognition algorithm with the result of recognizing an object using the second recognition algorithm (S2650). Step S2650 comprehensively illustrates the comparison process (matching) performed by the apparatus for improving object recognition rates, and may correspond to steps S2340 and S2350 in FIG. 23.

[0256] In one embodiment, the object recognition rate improving device receives input regarding a first reference value and a second reference value, and when at least one of the preset first reference value and second reference value is changed by the received first reference value or second reference value, the object recognition rate improving device may update the comparison result by referring to history information regarding frames before and after the already registered track. For example, when the first reference value is changed, the object recognition rate improving device may update the comparison result by referring to history information regarding a frame immediately before the track is generated, as described above with reference to FIG. 24. Furthermore, when the second reference value is changed, the object recognition rate improving device may update the comparison result by referring to history information regarding frames after the track is generated and deleted, as described above with reference to FIG. 25.

[0257] The apparatus for improving object recognition rates may correct the results of recognizing the object in the image using the first and second recognition algorithms based on the comparison results (S2670). Step S2670 comprehensively illustrates subsequent processing steps performed by the apparatus for improving object recognition rates, and may correspond to steps S2360, S2370, and S2380 in FIG. 23.

[0258] FIG. 27 is a block diagram of an object recognition rate improving device according to an embodiment.

[0259] 27, an object recognition rate improving device 2700 may include a communication unit 2710, a processor 2720, and a DB 2730. Only components related to the embodiment are shown in the object recognition rate improving device 2700 in Fig. 27. Therefore, a person skilled in the art would understand that the object recognition rate improving device 2700 may further include other general-purpose components in addition to the components shown in Fig. 27.

[0260] The communication unit 2710 may include one or more components that enable wired / wireless communication with an external server or device. For example, the communication unit 2710 may include at least one of a short-range communication unit (not shown), a mobile communication unit (not shown), and a broadcast receiving unit (not shown).

[0261] The DB 2730 is hardware that stores various data processed within the object recognition rate improving device 2700, and can store programs for processing and controlling the processor 2720.

[0262] DB2730 may include random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray or other optical disc storage, hard disk drive (HDD), solid state drive (SSD), or flash memory.

[0263] The processor 2720 controls the overall operation of the object recognition rate improving device 2700. For example, the processor 2720 may generally control an input unit (not shown), a display (not shown), a communication unit 2710, a DB 2730, etc. by executing a program stored in a DB 2730. The processor 2720 may control the operation of the object recognition rate improving device 2700 by executing a program stored in a DB 2730.

[0264] The processor 2720 may control at least a portion of the operations of the object recognition rate improving device 2700 described above.

[0265] As an example, as described with reference to Figures 23 to 26, the processor 2720 may recognize an object included in an image using a first recognition algorithm that recognizes an object for each frame of the image, form a track using multiple frames included in the image, recognize the object using a second recognition algorithm that recognizes the object included in the track, compare the result of recognizing the object using the first recognition algorithm with the result of recognizing the object using the second recognition algorithm, and correct the result of recognizing the object in the image using the first recognition algorithm and the second recognition algorithm based on the comparison result.

[0266] The processor 2720 may be implemented using at least one of ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.

[0267] The above-described embodiments of the present invention may be realized in the form of a computer program executable by various components on a computer, and such a computer program may be recorded on a computer-readable medium, including magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory.

[0268] On the other hand, the computer program may be one specially designed and constructed for the present invention, or it may be one that is well known and available to those skilled in the art of computer software. Examples of computer programs include not only machine language code such as that produced by a compiler, but also high-level language code that is executed by a computer using an interpreter, etc.

[0269] The specific implementation described in the present invention is merely an embodiment and does not limit the scope of the present invention in any way. For the sake of brevity, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of the system may be omitted. Furthermore, wire connections or connecting members between components shown in the drawings are illustrative of functional and / or physical or circuit connections, and may be replaced in actual devices and implemented as various additional functional, physical, or circuit connections. Furthermore, unless specifically referred to as "essential," "critically," or the like, a component may not necessarily be required for application of the present invention.

[0270] In the present specification (particularly in the claims), the use of the term "said" and similar indicators may refer to either the singular or the plural. Furthermore, when a range is described in the present invention, it encompasses the invention to which each individual value within that range is applied (unless otherwise specified), as if each individual value comprising that range were described in the detailed description of the invention. Finally, with respect to steps constituting the method of the present invention, unless explicitly stated or stated to the contrary, the steps may be performed in any suitable order. The present invention is not necessarily limited to the order of the steps described above. The use of all examples or exemplary terms (such as, for example, etc.) in the present invention is merely for the purpose of describing the present invention in detail, and the scope of the present invention is not limited by these examples or exemplary terms unless otherwise limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and variations can be made within the scope of the claims or their equivalents, depending on design conditions and factors.

Claims

1. applying a first recognition technique and a second recognition technique using different algorithms to a first image captured while driving to recognize an object included in the first image; applying a first detection technique and a second detection technique to the object recognition result, respectively, to detect frames for each of the applied detection techniques; aggregating the detected frames to generate a frame set including a plurality of frames; and sampling the merged set of frames to generate a second image; the first detection technique is a technique for detecting a specific frame in which an object is recognized from each of the recognition results obtained by the first recognition technique and the second recognition technique; A method for generating improved training data video, wherein the second detection technique is a technique for detecting a specific frame in which an object is re-recognized after disappearing from the recognition results by at least one of the first recognition technique and the second recognition technique.

2. The step of generating the second image includes: generating a frame group based on the integrated frame set, the frame group including at least one frame and the frames not overlapping with each other; The method of claim 1 , further comprising: extracting frames for each of the frame groups to generate the second image.

3. The step of generating the second image includes: The method of claim 2 , further comprising extracting one frame from each of the frame groups to generate the second image.

4. The step of generating the second image includes: The method of claim 2, further comprising extracting a number of frames corresponding to a weight set for each frame group to generate the second image.

5. The weight set for each frame group is The method for generating improved training data images according to claim 4 , wherein the value is determined based on the number of frames included in each frame group.

6. In the step of generating the second image, The method for generating an improved training data image according to claim 1, further comprising: sampling frames included in the integrated frame set based on a predetermined time interval to extract a plurality of frames; and generating the second image using the extracted frames.

7. In the step of generating the integrated frameset, Identifying frames detected as overlapping frames among the frames detected by the detection techniques; In the step of generating the second image, The method for generating improved training data images according to claim 1 , wherein the second image is generated by necessarily including the overlapping detected frames.

8. The first recognition technique is an algorithm for recognizing an object in the first image based on YoloV4-CSP; The second recognition technique is The method for generating improved training data images according to claim 1, wherein the algorithm for recognizing objects in the first image is based on YoloV4-P7.

9. The first detection technique comprises:

2. The method of claim 1, wherein the method is a detection technique that detects a specific frame in which the object is recognized based on a result of comparing frames of the object recognized by the first recognition technique and the second recognition technique.

10. The second detection technique comprises: The method for generating improved training data images according to claim 1, wherein the detection technique detects a specific frame in which an object recognized from the first image disappears and then reappears after a predetermined period of time, based on the result of detecting that the object has disappeared and then reappeared.

11. A computer-readable recording medium storing a program for executing the method of claim 1.

12. a memory having at least one program stored therein; a processor that performs calculations by executing the at least one program; The processor: applying a first recognition technique and a second recognition technique using different algorithms to a first image acquired while driving to recognize an object included in the first image; applying a first detection technique and a second detection technique to the object recognition result, respectively, to detect frames for each of the applied detection techniques; aggregating the detected frames to generate a frame set including a plurality of frames; Sampling the merged set of frames to generate a second image; the first detection technique is a technique for detecting a specific frame in which an object is recognized from each of the recognition results obtained by the first recognition technique and the second recognition technique; An apparatus for generating improved training data images, wherein the second detection technique is a technique for detecting a specific frame in which an object is re-recognized after disappearing from the recognition results by at least one of the first recognition technique and the second recognition technique.

Citation Information

Patent Citations

  • Training of an object recognition neural network

    CN113935395A

  • Image selection device and image selection method

    JP2020067818A

  • Vehicular illumination control system

    JP2020181310A

  • Training of object recognition neural network

    JP2022008187A

  • Method and apparatus for determining a driving route of a vehicle

    KR102438114B1