Method and system for estimating CCTV camera posture and object coordinates based on precise road map

The automated CCTV camera posture and object coordinate estimation using a precise road map addresses the challenges of rough estimation and low accuracy by automating the process, enhancing accuracy and efficiency in CCTV systems.

JP7770656B2Active Publication Date: 2025-11-17チェイング +3
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023207155
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-12-07
Publication Date
2025-11-17
Estimated Expiration
2043-12-07

AI Technical Summary

Technical Problem

Existing CCTV systems struggle with rough position estimation using the naked eye, difficulty in estimating global coordinates, and significantly low accuracy, making it challenging to share dangerous situation information accurately and efficiently.

Method used

An automated method for estimating CCTV camera posture and object coordinates using a precise road map, which includes matching road objects in the camera's image with a precise road map, selecting mapping targets, and correcting errors to enhance accuracy.

Benefits of technology

This method automates coordinate estimation and mapping, achieving higher accuracy and reducing time and costs without on-site work, applicable to both fixed and PTZ cameras, and enables accurate sharing of danger information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007770656000001
    Figure 0007770656000001
  • Figure 0007770656000002
    Figure 0007770656000002
  • Figure 0007770656000003
    Figure 0007770656000003
Patent Text Reader

Abstract

To provide an automatic position coordinate estimation technique for an object or area in an image captured by a CCTV camera and an automatic mapping technique using the same to solve problems such as rough position estimation by the naked eye, difficulty in global coordinate estimation, and remarkably low accuracy even when estimated.SOLUTION: A CCTV camera attitude and object coordinate estimation method for a precision road map base, includes the steps for: (a) estimating a camera attitude by matching a road surface object in a captured video of the CCTV camera with a precision road map; (b) selecting an object to be mapped in the captured video; and (c) estimating coordinates of the selected object to be mapped in the captured video based on the estimated camera attitude. The camera attitude includes information for the position (X, Y, and Z), pan, and tilt of the CCTV camera.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for estimating CCTV camera pose and object coordinates based on a precise road map. [Background technology]

[0002] Currently, thousands of CCTVs are installed along expressways to monitor traffic conditions, and even in urban areas, CCTVs are installed at many intersections and used for traffic monitoring.In addition, with the recent development of AI technology, it is now possible to recognize vehicles and people, and by using this, it is possible to recognize stopped vehicles, pedestrians, and vehicles driving in the wrong direction.

[0003] However, even when dangerous situations are automatically recognized, the operator must roughly estimate the location through visual judgment, or in most cases the location cannot be estimated using a global coordinate system of latitude, longitude, and altitude. Even if an estimate is made, there are large errors, making it difficult to share information.

[0004] In particular, up until now, most CCTV systems have required operators to recognize dangerous situations through visual monitoring and then share the information after estimating the approximate location through visual judgment. This takes too much time to generate the information and significantly reduces accuracy, making it difficult to share the information through mapping. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Korean Patent Publication No. 10-2021-0050997 Summary of the Invention [Problem to be solved by the invention]

[0006] In order to solve problems such as rough position estimation using the naked eye, difficulty in estimating global coordinates, and significantly low accuracy even if estimation is performed, the present invention aims to provide an automated position coordinate estimation technology for objects and areas in images captured by CCTV cameras and a technology related to automatic mapping using the same.

[0007] In addition to fixed cameras, we will develop technology that can be applied to PTZ cameras that rotate at any time, thereby maximizing the range of applications.

[0008] In particular, we propose a coordinate estimation and mapping technology that automatically estimates the camera orientation of CCTV cameras using the HD Map, which is currently being constructed by the government based on major roads, and uses the estimated camera orientation. In particular, we provide a method to automatically correct position errors and camera orientation by checking the error in the coordinate estimation result and using the result.

[0009] Therefore, once the construction of precise road maps is completed nationwide, the present invention can be applied to automatically estimate the precise location of objects and dangerous areas recognized by AI, automatically map them, and share danger information. [Means for solving the problem]

[0010] To solve the above problem, a method for estimating CCTV camera posture and object coordinates based on a precise road map according to one embodiment of the present invention includes: (a) estimating a camera posture by matching road objects in a CCTV camera image with a precise road map; (b) selecting a mapping target object in the image; and (c) estimating the coordinates of the mapping target object selected in the image based on the estimated camera posture, wherein the camera posture includes information on the position (X, Y, Z), pan, and tilt of the CCTV camera.

[0011] Step (a) may include extracting pixel coordinates of feature points of the road surface object from the captured image; matching the road surface object with a precise road map to extract spatial coordinates of the feature points of the road surface object; and estimating the camera pose using Perspective-n-Point (PnP) from the relationship between the pixel coordinates and spatial coordinates of the extracted feature points of the road surface object.

[0012] If the CCTV camera is a PTZ camera, step (a) may further include converting the pan and tilt values ​​provided by the PTZ camera into absolute values ​​based on true north and a vertical direction based on the difference between the pan and tilt values ​​of the estimated camera attitude and the pan and tilt values ​​provided by the PTZ camera; and reducing the captured image by the reciprocal of the zoom value provided by the PTZ camera to convert it into an image with a zoom value of “0”.

[0013] The step (a) may further include: (a1) projecting the precise road map onto the photographed image using the estimated camera pose, or back-projecting an object in the photographed image onto the precise road map using the estimated camera pose; (a2) calculating a pixel error of each feature point of the corresponding object in the photographed image and the object in the precise road map; and (a3) ​​correcting the estimated camera pose to reduce the calculated pixel error below the critical value if the calculated pixel error is equal to or greater than a critical value.

[0014] In step (b), the mapping target object may be an area or object that is automatically recognized and extracted from the captured image according to a predetermined rule based on AI (artificial intelligence), or an area or object that is manually selected from the captured image by a user.

[0015] Step (b) may include back-projecting each feature point of the mapping target object onto the precise road map according to the estimated camera pose; and calculating spatial coordinates of each back-projected feature point of the mapping target object from the precise road map. [Effects of the Invention]

[0016] Therefore, the CCTV camera pose and object coordinate estimation method and system according to one embodiment of the present invention provides a method for estimating the CCTV camera pose through a simple method using a precision road map without on-site work such as surveying or setting control points, thereby automating coordinate estimation and mapping, and achieving effects such as higher accuracy and time and cost savings compared to existing methods.

[0017] It can be applied to not only fixed CCTV but also PTZ cameras, and can automate the coordinate estimation and mapping process. It can easily improve accuracy by estimating coordinates on a precise road map, and can be applied not only when it is possible to recognize road objects such as lanes using AI, but also when such a function is not available.

[0018] In addition, errors in the coordinate estimation results and camera attitude are all checked and automatically corrected during the estimation process, enabling highly accurate estimation. If road surface objects can be recognized, the entire process can be automated, increasing convenience. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a network configuration diagram of a method and system for estimating CCTV camera posture and object coordinates based on a precise road map according to an embodiment of the present invention. [Figure 2] 2 is a detailed configuration diagram of the information processing server shown in FIG. 1; [Figure 3] 1 is a flowchart illustrating a method for estimating CCTV camera posture and object coordinates based on a precise road map according to an embodiment of the present invention. [Figure 4] 4 is a detailed flowchart of the camera pose estimation step shown in FIG. 3. [Figure 5] 10 is an exemplary diagram illustrating a fine adjustment procedure after projecting a precise road map onto a captured image. FIG. [Figure 6] 4 is a flowchart illustrating steps from a coordinate estimation step to a message generation and distribution step shown in FIG. 3. [Figure 7] 10 is an exemplary diagram illustrating the concept of attitude correction and mapping error correction through lane designation adjacent to a dangerous area; FIG. DETAILED DESCRIPTION OF THE INVENTION

[0020] The description of the present invention is merely an example for structural or functional explanation, and therefore the scope of the present invention should not be construed as being limited by the examples described herein. In other words, since the examples can be modified in various ways and can have various forms, the scope of the present invention includes equivalents that can embody the technical idea. Furthermore, the objectives or effects presented in the present invention do not mean that a particular example must include all of these or only such effects, and therefore the scope of the present invention should not be construed as being limited thereby.

[0021] Meanwhile, the meanings of the terms described in this invention shall be understood as follows. Terms such as "first" and "second" are used to distinguish one component from another, and shall not limit the scope of rights. For example, a first component may be named a second component, and similarly, a second component may be named a first component. When a component is referred to as being "connected" to another component, it is understood that it may be directly connected to the other component, but that there may be other components between them. Conversely, when a component is referred to as being "directly connected" to another component, it is understood that there are no other components between them. Meanwhile, other expressions describing the relationship between components, such as "between" and "immediately between," or "adjacent to" and "directly adjacent to," shall be interpreted similarly. Singular expressions should be understood to include plural expressions unless the context clearly dictates otherwise, and terms such as "comprise" or "have" are intended to specify the presence of stated features, numbers, steps, operations, components, parts, or combinations thereof, but are understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0022] Hereinafter, a system and method for estimating roadside CCTV postures, estimating AI-recognized object coordinates, and automating dangerous area mapping based on a precise road map according to an embodiment of the present invention will be described in more detail with reference to the accompanying drawings.

[0023] 1 is a network configuration diagram of a system for estimating CCTV camera posture and object coordinates based on a precise road map according to an embodiment of the present invention, and FIG. 2 is a detailed configuration diagram of an information processing server shown in FIG.

[0024] First, as shown in FIG. 1, a precise road map-based CCTV camera pose and object coordinate estimation system 100 according to one embodiment of the present invention includes a CCTV camera 10 and an information processing server 20, and may further include an operator terminal 30.

[0025] If each component communicates over a network, the network means a connection structure that allows information exchange between each node, such as a plurality of terminals and servers. Examples of such networks include, but are not limited to, RF, 3GPP (registered trademark) (3rd Generation Partnership Project) network, LTE (Long Term Evolution) network, 5GPP (5th Generation Partnership Project) network, WIMAX (World Interoperability for Microwave Access) network, Internet, LAN (Local Area Network), WirelessLAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth (registered trademark) network, NFC network, satellite broadcasting network, analog broadcasting network, DMB (Digital Multimedia Broadcasting) network, etc.

[0026] In the following, the term "at least one" is defined as a term including both singular and plural, and it is clear that even if the term "at least one" is not present, each component can exist in singular or plural and can mean singular or plural. Furthermore, it can be said that whether each component is provided in singular or plural can be changed depending on the embodiment.

[0027] The CCTV camera 10 may be a PTZ (pan tilt zoom) camera or a fixed camera installed on the side of a road. If the CCTV camera is a PTZ camera, it may be a device that provides information on the camera's position (X, Y, Z), pan, tilt, and zoom.

[0028] Next, the information processing server 20 estimates the camera pose (X, Y, Z, Pan, Tilt, Zoom) of the CCTV camera using the video captured by the CCTV camera 10 and a high definition map (HD Map). If the CCTV camera is fixed, the estimated camera pose may be used until an error exceeding a critical value is detected. If the CCTV camera is a PTZ camera, the camera pose may be estimated each time the pan, tilt, or zoom is changed.

[0029] In addition, the information processing server 20 sets as mapping target objects objects or areas that are automatically recognized and extracted from the captured image by the CCTV camera according to pre-set rules based on AI (artificial intelligence), or areas or objects that are manually selected by the user from the captured image.

[0030] The objects to be mapped may be moving objects or dangerous areas that meet the rules based on AI (Artificial Intelligence), or objects or dangerous areas that the operator recognizes with the naked eye may be selected as the objects to be mapped.

[0031] The information processing server 20 may be configured to automatically estimate the coordinates of characteristic points of mapping targets, such as moving objects and dangerous areas, that are automatically recognized by AI or selected by an operator, based on a precise road map, and in the process, correct pre-estimated position errors and camera attitudes using the precise road map, and then generate a message containing surface-based information on types of danger and prohibited areas, and distribute it using V2X, mobile communication, etc.

[0032] Next, the operator terminal 30 may be configured to provide the information processing server 20 with setting information that sets the danger zone within the precise road map.

[0033] More specifically, the information processing server 20 may include at least one of a camera posture estimation unit 21 , a mapping target object setting unit 22 , a coordinate estimation unit 23 , a correction unit 24 , and a message generation and distribution unit 25 .

[0034] The camera posture estimation unit 21 matches road objects in the image captured by the CCTV camera with corresponding objects in the precise road map to estimate the camera posture. Here, the camera posture includes information on the position (X, Y, Z), pan, and tilt of the CCTV camera, and may further include information on zoom.

[0035] The camera posture estimation unit 21 extracts pixel coordinates of feature points of road surface objects from the captured image, matches the road surface objects with a precise road map, extracts spatial coordinates of each feature point of the road surface objects, and estimates the camera posture using PnP (Perspective-n-Point) from the relationship between the pixel coordinates and spatial coordinates of each extracted feature point of the road surface objects.

[0036] PnP (Perspective-n-Point) is a method of estimating the camera posture from the positional relationship of three or more 3D points projected onto a 2D plane. In this invention, three or more feature points are extracted from the road surface object, and the camera posture can be estimated using PnP (Perspective-n-Point) from the relationship between the pixel coordinates of the feature points and the spatial coordinates.

[0037] Meanwhile, the feature point in the present invention refers to a point that can show a feature that can distinguish an object present in an image, and can be a corner point or a point that can be distinguished from its surroundings by color or brightness.

[0038] If the CCTV camera is fixed, the camera position (X, Y, Z), pan, tilt, and zoom do not change, so the initially estimated camera posture can continue to be used. Of course, even in the case of a fixed camera, the camera may shake for various reasons, or there may be an error in the initially estimated camera posture, so the correction unit 24, described below, performs correction on the camera posture through various processes.

[0039] If the CCTV camera is a PTZ camera, the initial estimated camera posture can continue to be used because the posture can be maintained unless pan, tilt, or zoom is performed.If at least one of pan, tilt, and zoom operations is performed, the corresponding operation value can be obtained, so the camera posture estimation unit 21 can obtain a modified camera posture by considering the PTZ (PanTiltZoom) value provided by the camera to the initial camera posture.Since the camera installation position is fixed, the camera position (X, Y, Z) continues to be maintained unless correction is required.

[0040] Meanwhile, if the CCTV camera is a PTZ camera, the camera posture estimation unit 21 can easily obtain the PT value from the difference between the pan and tilt values ​​of the estimated camera posture and the pan and tilt values ​​provided by the PTZ camera by converting the pan and tilt values ​​provided by the PTZ camera into absolute values ​​based on true north and the vertical direction. Also, if the CCTV camera is a PTZ camera, the camera posture estimation unit 21 can perform a process of reducing the captured image by the reciprocal of the zoom value provided by the PTZ camera to convert it into a captured image with a zoom value of '0'.

[0041] Meanwhile, the correction unit 24 can correct information regarding the camera pose and / or information regarding the coordinates of the object to be mapped.

[0042] The correction unit 24 projects a precise road map onto the photographed image using the estimated camera attitude, calculates pixel errors of objects in the photographed image and each feature point of the corresponding object in the precise road map, and if the calculated pixel errors are greater than a critical value, corrects the camera attitude and repeats the above process to reduce the pixel errors to less than the critical value.

[0043] Of course, the correction unit 24 can correct the attitude by reversing the above process, but it can also back-project the object in the captured image onto the precise road map using the estimated camera attitude, calculate the position error (coordinate error) of each feature point of the object in the back-projected captured image and the corresponding object in the precise road map, and if the calculated position error is greater than a critical value, correct the camera attitude and repeat the above process to reduce the pixel error to less than the critical value.

[0044] The correction unit 24 calculates at least one of pixel error and position error (coordinate error), and repeats the process of finely correcting the camera attitude to reduce the corresponding error below a critical value, thereby minimizing the camera attitude and position error.

[0045] In the present invention as a whole, the correction process of the correction unit 24 may be one of the following: i) a method of projecting a precise road map onto a photographed image using an estimated camera attitude, and then correcting using pixel errors between objects on the precise road map and corresponding image objects, and ii) a method of back-projecting a photographed image onto a precise road map using an estimated camera attitude, and then correcting using coordinate errors between objects on the precise road map and corresponding image objects. The correction processes i) and ii) are inversely related, and the correction methods and results are the same, so either one of the two methods or both of them may be used.

[0046] The mapping target object setting unit 22 sets, as a mapping target object, an object or area that is automatically recognized and extracted in the photographed image according to a predetermined rule based on AI (artificial intelligence), or an object or area that is manually selected in the photographed image by a user (operator).

[0047] The mapping target object setting unit 22 may be configured to automatically estimate the coordinates of objects such as vehicles and people recognized based on AI such as deep learning, in addition to road surface objects such as lanes, based on a precise road map, or to automatically map automatically recognized hazardous elements or areas on the road occupied by the hazardous elements using AID (Automatic Incident Detection) as mapping target objects.

[0048] The operation of mapping a recognized object such as a vehicle or person or a dangerous area recognized by AID is as follows.

[0049] The system extracts pixel coordinates of feature points (three or more feature points) of road surface objects such as lanes, recognizes mapping target objects other than road surface objects (including AID-based hazardous areas), back-projects the mapping target objects and feature points of the road surface objects onto a precise road map, or calculates the position error between the back-projected road surface object and the corresponding object on the precise road map, calculates the spatial coordinates of the mapping target feature points using the precise road map, and corrects the spatial coordinates of the mapping target feature points.The system then re-projects the mapping target feature points onto an image, calculates the pixel error for each re-projected feature point based on the initially set point, and if the pixel error is smaller than a critical value, displays the mapping target (position, attributes) on the precise road map, and automatically generates and displays a message.

[0050] Here, if the pixel error is greater than or equal to a critical value, the pose (pan, tilt) can be corrected using the position error for each feature point.

[0051] In addition, the mapping target object setting unit 22 can set an object set or designated by an operator or a dangerous area as a mapping target object. To this end, the mapping target object setting unit 22 recognizes road surface objects on a precise road map, such as lanes and arrows, and extracts pixel coordinates of feature points (three or more feature points) of the recognized road surface objects.

[0052] When the operator inputs the settings and attributes of a dangerous area (three or more feature points) in the captured image, the coordinate estimation unit 23 described later can back-project the dangerous area and the feature points of the road surface object onto a precise road map and calculate the spatial coordinates of the feature points of the dangerous area using the precise road map.

[0053] Meanwhile, the correction unit 24 calculates the position error between the back-projected road surface object and the same object on the precise road map, corrects the spatial coordinates of the feature points of the dangerous area reflecting the position error, re-projects the feature points of the dangerous area onto the captured image, and calculates the pixel error for each re-projected feature point based on the minimum set point. After the calculation is completed, if the pixel error is smaller than a critical value, the message generation and distribution unit 25 may be configured to mark the dangerous area (position, attribute) on the precise road map and automatically generate and display a message.

[0054] Here, if the pixel error is equal to or greater than a critical value, the correction unit 24 can correct the posture (pan, tilt) using the position error for each feature point.

[0055] The coordinate estimation unit 23 can automatically estimate the coordinates of feature points for mapping target objects (objects recognized by AI, dangerous areas recognized by AID, objects or dangerous areas selected by an operator) based on a precise road map.

[0056] The correction unit 24 may be configured to correct at least one of pixel error, position error, and camera attitude that are estimated in advance using a precise road map in a camera attitude estimation process or a coordinate estimation process of a mapping target object.

[0057] The correction unit 24 may be configured to correct camera attitude and position errors in the process of estimating or determining objects recognized by AI, such as vehicles, people, and signs, danger areas recognized by AID, or danger areas set by an operator.

[0058] The correction unit 24 projects the precise road map onto the photographed image, and then automatically determines whether to perform correction according to a preset standard based on the position error between road surface objects such as lanes in the photographed image and road surface objects on the projected precise road map. At this time, the recognized road surface objects are back-projected onto the precise road map, or the same objects on the precise road map are projected onto the photographed image, and then the coordinates of the mapping target objects and dangerous areas are automatically corrected through position comparison (coordinate comparison), and the posture is also corrected.

[0059] Even in the case of a fixed type, there is a high possibility that rotation errors will occur due to displacement over time, so accuracy can be maintained through correction at each mapping. Once the initial camera posture estimation is complete, there is almost no possibility of the position (X, Y, Z) changing, so it is assumed that there is no error, and the error type can be classified as Pan and Tilt.

[0060] Left-right errors can be corrected by fine-tuning Pan, and up-down errors by fine-tuning Tilt. After calculating the magnitude of the error and determining the adjustment unit, the process can be repeated until an error in the opposite direction occurs, and then fine-tuning the opposite direction to 1 / 10 of the original unit, and then repeating the process until an error in the opposite direction occurs again. At each stage, a comparison with a threshold is performed, and if satisfied, the process can be terminated.

[0061] Meanwhile, if there is an error in the PTZ value when calculating coordinates on the precise road map after back-projecting the mapping target object or dangerous area designated by the operator using only the PTZ value provided by the camera, there is a possibility that the coordinate estimation result will have a large error.

[0062] The operator can check the magnitude of such errors by estimating the coordinates of the dangerous area on a precise road map, projecting it onto the photographed image together with surrounding road surface objects, and visually comparing the positions of the road surface objects on the precise road map projected onto the photographed image with the positions of the corresponding road surface objects in the existing photographed image. If the operator determines that a large error has occurred as a result of visually checking, the operator can mark the start and end points of the road surface objects in the photographed image that are the same as the projected road surface object, and automatically correct the camera posture using this, and then perform the coordinate estimation process again using this to further correct the error.

[0063] Even in the case of a fixed type, there is a high possibility that rotation errors will occur due to displacement over time, so accuracy can be maintained through correction at each mapping.

[0064] That is, once the initial camera posture estimation is complete, it is assumed that there is no error because the position (X, Y, Z) is unlikely to change, and the error types can be classified into Pan and Tilt.

[0065] Left-right errors can be corrected by fine-tuning Pan, and up-down errors by fine-tuning Tilt. The correction procedure involves calculating the magnitude of the error, then determining the adjustment unit → repeating until an error in the opposite direction occurs → fine-tuning in the opposite direction to 1 / 10 of the original unit → repeating until an error in the opposite direction occurs again. At each stage, a comparison with a threshold is made, and if satisfied, the process is terminated.

[0066] The message generation and distribution unit 26 may be configured to generate a message including the type of danger, surface-based information on prohibited areas, etc., and distribute it using V2X, mobile communication, etc.

[0067] Fig. 3 is a flowchart illustrating a method for estimating CCTV camera attitude and object coordinates based on a precise road map according to an embodiment of the present invention, Fig. 4 is a detailed flowchart of the camera attitude estimation step shown in Fig. 3, Fig. 5 is an example diagram illustrating a fine adjustment procedure after projecting a precise road map onto a captured image, and Fig. 6 is a flowchart illustrating the coordinate estimation step through the message generation and distribution step shown in Fig. 3. Fig. 7 is an example diagram illustrating the concept of attitude correction and mapping error correction through designation of lanes adjacent to a dangerous area.

[0068] First, as shown in FIG. 3, a method for estimating CCTV camera posture and object coordinates based on a precise road map according to one embodiment of the present invention (S700) includes at least one of a camera posture estimation step (S710), a mapping target object setting step (S720), a coordinate estimation step (S730), an error correction step (S740), and a message generation and distribution step (S750).

[0069] The camera attitude estimation step (S710) may be a step of estimating the camera attitude (X, Y, Z, Pan, Tilt) of the CCTV camera, determining the position (X, Y, Z) of the PTZ camera, and determining a relational equation (difference) for converting the Pan / Tilt value provided by the PTZ camera into an absolute value (true north direction, horizontal reference).

[0070] The camera pose estimation referred to in this application is a technology for estimating the pose of a CCTV camera through a simple process using a precision road map without on-site work such as surveying, setting control points, and marking, and can provide a foundation for automating coordinate estimation and mapping, which will be described later, and can also be a technology that provides effects such as high accuracy and time and cost savings compared to existing methods. Therefore, in the case of a PTZ camera, since the device provides relative values ​​based on the initial setting position, it is not possible to immediately convert them to a camera pose, so this can be a technology that converts PTZ values ​​provided thereafter and uses them as a camera pose by calculating and reflecting the difference between the value provided when estimating the initial camera pose and the true value.

[0071] Next, the step of setting the mapping target object (S720) may include at least one of AI-based mapping target object recognition, AID (Automatic Incident Detection)-based object or danger area recognition, and a step of designating an object or area that an operator recognizes as a dangerous situation with the naked eye as a danger area.

[0072] Next, the coordinate estimation step (S730) may be a step of automatically estimating and mapping the coordinates of feature points based on a precise road map for the object to be mapped (objects or areas recognized by AI, objects or dangerous areas recognized by AID, objects or dangerous areas manually set by an operator, etc.).

[0073] The coordinate estimation of the object to be mapped can be applied to cases where not only fixed CCTV but also PTZ cameras are used, and is a method that can automate the coordinate estimation and mapping process, and can easily and accurately estimate coordinates on a precise road map. In particular, it can be applied not only when it is possible to recognize road surface objects such as lanes using AI (Artificial Intelligence), but also when such a function is not available.

[0074] Next, the error correction step (S740) may be a step of correcting the position, pixel error, and camera attitude previously estimated using the precise road map in the camera attitude estimation step (S710) or the coordinate estimation step (S730) of the mapping target object.

[0075] The position, pixel error, and / or attitude correction step is a method that enables highly accurate estimation by checking and automatically correcting errors in the coordinate estimation results and the camera attitude applied at this time throughout the entire estimation process. In particular, when it is possible to recognize road objects such as lanes, the entire process can be automated, thereby increasing convenience. Even if such a function is not available, this technology can achieve the same effect with just a simple operation by an operator.

[0076] Next, the message generation and distribution step (S750) may be a step of automatically generating and distributing risk-related messages.

[0077] In sharing information about the mapping target object created based on the above process, a message is automatically generated so that no additional manual work is required by the operator, and the result is displayed on a monitor managed by the operator so that the operator can check it, and distribution can also be done automatically.

[0078] Each of the above steps will now be described in more detail.

[0079] 4, the initial camera posture estimation step (S710) extracts pixel coordinates of feature points of road objects in CCTV images, matches them with a precise road map to extract spatial coordinates of the feature points of the road objects, estimates the camera posture through Perspective-n-Point (PnP), and then projects the precise road map onto the captured image using the estimated camera posture. Thereafter, pixel errors within the image for each feature point are calculated, and the pixel errors are compared with a threshold.

[0080] As mentioned above, PnP (Perspective-n-Point) refers to a technique for obtaining the accurate position and orientation of a camera that captures an image in a 3D coordinate system based on the coordinates of 3D feature points whose coordinate positions are known and the coordinates of 2D feature points on an image onto which the corresponding feature point coordinates are projected.

[0081] If the error is below the threshold, the process ends. If the error is above the threshold, the camera posture is finely corrected based on the error characteristics, and a precise road map is projected onto the captured image using the corrected camera posture. The process can then be repeated until the pixel error in the image for each feature point falls below the threshold.

[0082] Meanwhile, the camera pose estimation step (S710) may be performed as follows after estimating the primary pose through PnP.

[0083] For example, the method may include back-projecting the feature points of the road surface object onto a precise road map based on the estimated camera posture, calculating a position error between the spatial coordinates of the back-projected feature points of the road surface object and the corresponding spatial coordinates of the object on the precise road map for each feature point, and terminating the process if the error is less than a threshold, and fine-correcting the estimated camera posture in consideration of the error characteristics if the error is greater than or equal to the threshold, and then back-projecting the road surface object onto the precise road map using the corrected posture. Again, the process may be repeated until the position error in the space for each feature point becomes less than the threshold.

[0084] Here, in the PnP-based primary posture estimation, in the case of a rotating PTZ camera, it is efficient to set the camera approximately perpendicular to the road.

[0085] Based on roughly known CCTV camera position information and information on road surface or roadside objects (shown on a precise road map), the precise road map is compared to match road surface objects such as lanes, and using the rough position information and surrounding objects, matching pairs of objects on the image and the precise road map are searched for with the naked eye using the rough position information and surrounding objects (surrounding objects: traffic lights, signs, road markings other than lanes, start and end points of bridges, start and end points of protective facilities, etc.), pixel coordinates and spatial coordinates for three or more matching pairs are extracted, and then PnP can be performed.

[0086] In addition, in the process of finely correcting the posture in consideration of the error characteristics, the process may involve estimating the primary camera posture based on PnP, projecting a precise road map onto the captured image using this camera posture, calculating pixel errors between object feature points to confirm the degree of error, and finally correcting the camera posture through fine adjustments that reflect the error characteristics.

[0087] Alternatively, the camera posture may be used to back-project feature points of specific objects in a captured image onto a precise road map, and then correct the camera posture based on the position error between the feature points of the corresponding objects.

[0088] In the case of a PTZ camera, the camera posture estimation step (S710) performs camera calibration to determine true values, and then converts the difference between the provided Pan and Tilt values ​​and the true values ​​into absolute values. In the case of zoom, the magnification ratio is calculated using the provided values, and then the image is reduced by the reciprocal of this ratio, so that the image is converted into an image with a zoom value of '0'. The step can also include a process of continuously correcting the difference using the corrected camera posture value in an automated mapping process.

[0089] Meanwhile, the mapping target object setting step (S720) may have different processes depending on the method of setting recognized objects such as vehicles and people recognized based on AI, objects and danger areas recognized by AID, and objects and danger areas set by an operator.

[0090] First, referring to FIG. 6, pixel coordinates of characteristic points of road surface objects such as lanes are extracted, and mapping targets (including AID-based danger zones) other than road surface objects are recognized.

[0091] Then, the feature points of the object to be mapped and the road surface object are back-projected onto the precise road map to calculate the spatial coordinates of the feature points of the object to be mapped, the position error between the back-projected object and the same object on the precise road map is calculated, and the spatial coordinates of the feature points of the object to be mapped are corrected. Then, the feature points of the object to be mapped are re-projected onto the image, and a pixel error for each re-projected feature point is calculated based on the initially set point, and if the pixel error is smaller than a threshold, the method can include a process of notating the object to be mapped (position, attribute) on the precise road map and automatically generating and displaying a message.

[0092] Here, if the pixel error is greater than or equal to a critical value, the method may further include a process of correcting the camera pose (pan, tilt) using the position error for each feature point.

[0093] When an object such as a vehicle, person, fallen object, or animal is recognized, the coordinates of the object are automatically estimated. When a dangerous area such as an accident, wrong-way driving, or congestion is automatically recognized using AID (Automatic Incident Detection), the coordinates of the dangerous area are automatically estimated and the results are reflected in the automatic mapping.

[0094] Here, AI-based recognition targets include objects requiring coordinate estimation, moving objects such as vehicles, people, and animals, atypical obstacles such as wrong-way vehicles and fallen objects, the tail of traffic jams, traffic cones (registered trademark), and construction-related objects such as precast barriers.

[0095] The danger zone requiring coordinate estimation may be an accident zone recognized by AID, a construction zone recognized based on a series of traffic cones or precast barriers, or other danger zones recognized by surface.

[0096] In addition, the road surface objects for error confirmation and attitude correction may be road surface objects included in a precise road map, such as lanes, arrows, and stop lines (including signboards and traffic lights, if necessary).

[0097] Here, in the case of a PTZ camera, the initial camera posture estimation state may be a state in which the absolute values ​​(true values) of Pan and Tilt can be estimated in real time even during rotation, since the difference between the provided value and the true value is known.

[0098] In addition, after projecting a precise road map onto the captured image, the system automatically determines whether to correct position and orientation errors based on a preset standard, based on the position error between road objects in the captured image and those on the projected precise road map. This system also includes a function to back-project road objects recognized in the image onto the precise road map, or to project the same objects on the precise road map onto the captured image, and then automatically correct the coordinates of mapping objects and dangerous areas and correct their orientations by comparing the positions of the corresponding objects.

[0099] Even in the case of a fixed system, there is a high possibility that rotation errors will occur due to displacement over time, so accuracy can be maintained through corrections at each mapping.Once the initial camera posture estimation is complete, there is almost no possibility of position (X, Y, Z) changes, so it is assumed that there are no errors, but the types of errors can be classified as Pan and Tilt.

[0100] Left-right errors can be corrected by fine-tuning Pan, and up-down errors by fine-tuning Tilt. The correction procedure calculates the magnitude of the error and determines the adjustment unit, then repeats until an error in the opposite direction occurs, fine-tuning it to 1 / 10 of the original unit in the opposite direction, and repeats again until an error in the opposite direction occurs. At each stage, a comparison with a threshold is made and the process ends if the result is satisfactory.

[0101] When an object such as a vehicle, person, fallen object, or animal is recognized, the coordinates of the object are automatically estimated. When a dangerous area such as an accident, wrong-way driving, or congestion is automatically recognized using AID (Automatic Incident Detection), the coordinates of the dangerous area are automatically estimated and the results are reflected in the automatic mapping.

[0102] AI-based recognition targets are objects that require coordinate estimation, and include moving objects such as vehicles, people, and animals (including atypical obstacles such as wrong-way vehicles and fallen objects, the tail of traffic jams, traffic cones, precast barriers, and other construction-related objects).

[0103] In addition, the danger zone requiring coordinate estimation may be a danger zone recognized by a surface such as an accident zone recognized by AID, or a construction zone recognized based on a series of traffic cones or precast barriers.

[0104] On the other hand, if the camera is a PTZ camera, the difference between the provided value and the true value is known during the camera posture estimation process, so it may be possible to estimate the absolute values ​​(true values) of Pan and Tilt in real time even during rotation.

[0105] In addition, during the error correction process, when the mapping object or danger zone designated by the operator is back-projected onto the precise road map using only the provided PTZ value and then the coordinates are calculated on the precise road map, if there is an error in the PT value, there is a possibility that a large error will occur in the coordinate estimation results.

[0106] Such errors can be corrected by estimating the coordinates of the danger zone on a precise road map, re-projecting the surrounding road surface objects onto the captured image using the estimated camera attitude, and then having the operator visually compare the position of the re-projected road surface object in the image with the position of the corresponding same road surface object in the existing captured image. If the operator determines that the error is large after visually checking, they can mark the start and end points of the road surface object in the captured image that correspond to the re-projected road surface object, and automatically correct the camera attitude using this, and then performing the coordinate estimation process again using this to correct the coordinate estimation error.

[0107] Even in the case of a fixed type, there is a high possibility that rotation errors will occur due to displacement over time, so accuracy can be maintained through correction at each mapping.

[0108] That is, once the initial camera posture estimation is complete, the position (X, Y, Z) is unlikely to change, so it is assumed that there is no error, and the error types can be classified into Pan and Tilt.

[0109] Left and right errors can be corrected by fine-tuning Pan, and up and down errors by fine-tuning Tilt. The correction procedure is to calculate the magnitude of the error and determine the adjustment unit → repeat until an error in the opposite direction occurs → fine-tune in the opposite direction to 1 / 10 of the first unit → repeat until an error in the opposite direction occurs again. At each stage, a comparison with a threshold is performed and if satisfied, the process is terminated.

[0110] Meanwhile, an operator can set a danger zone (three or more feature points) within the captured image or input attributes (operator). In this case, the start and end points of adjacent road objects adjacent to the danger zone are set, and the feature points of the danger zone and adjacent road objects are back-projected onto a precise road map. The spatial coordinates of the feature points of the danger zone and adjacent road objects can then be calculated using the precise road map. At this time, the position error between the back-projected adjacent road objects and their corresponding objects on the precise road map can also be calculated. Then, the spatial coordinates of the feature points of the danger zone are corrected using the position error, and the feature points of the danger zone are re-projected onto the captured image. The pixel error for each re-projected feature point is calculated based on the initially set point, and if the error is smaller than a critical value, the system proceeds to the process of notating the danger zone (location, attributes) of the mapping target on the precise road map and automatically generating and displaying a message.

[0111] Here, if the error is greater than a critical value, the method may further include a process of correcting the camera posture (pan, tilt) using the position error for each feature point.

[0112] The method for estimating CCTV camera posture and object coordinates based on a precise road map according to one embodiment of the present invention provides a method for estimating the posture of a CCTV camera through a simple method using a precise road map without on-site work such as surveying, setting control points, and marking, thereby providing a foundation for automating coordinate estimation and mapping, and also has the advantages of producing effects such as high accuracy and time and cost savings compared to existing methods.

[0113] In addition, in the case of a PTZ camera, since the device provides relative values ​​based on the initial setting position, it is not possible to immediately convert them to the camera posture. Therefore, by calculating and reflecting the difference between the value provided at the time of initial posture estimation and the true value, it is possible to convert the PTZ values ​​provided thereafter and immediately use them as the camera posture.

[0114] In addition, it can be applied to not only fixed CCTV but also PTZ cameras, and presents a method to automate the coordinate estimation and mapping process, estimating coordinates on a precise road map is simple and improves accuracy. In particular, it has the advantage of being applicable not only when it is possible to recognize road objects such as lanes using AI, but also when such a function is not available.

[0115] Another advantage is that it is possible to automate error and attitude correction, for example, by checking and automatically correcting errors in the coordinate estimation results and the camera attitude applied at that time throughout the entire estimation process, highly accurate estimation is possible, and in particular, when it is possible to recognize road objects such as lanes, the entire process can be automated, thereby increasing convenience.Even if such a function is not available, it is possible to implement an algorithm that can achieve the same effect with just a simple operation by the operator.

[0116] Another advantage is that messages can be automatically generated and distributed. For example, messages can be generated automatically so that no additional manual work is required by the operator when sharing mapping information created based on the above process. The results can be displayed on a monitor managed by the operator so that the operator can check them, and distribution can also be done automatically.

[0117] Meanwhile, the components disclosed herein may be implemented using one or more general-purpose or special-purpose computers, including, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of software. For ease of understanding, a single processing device may be described; however, those skilled in the art will recognize that a processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing device may include multiple processors or one processor and one controller. Other processing configurations, such as parallel processors, are also possible.

[0118] Software may include a computer program, code, instructions, or a combination of one or more of these, which can configure a processing device to operate as desired or independently or collectively instruct the processing device. The software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave to be interpreted by the processing device or to provide instructions or data to the processing device. The software can also be distributed across computer systems coupled to a network, stored and executed in a distributed manner. The software and data can be stored on one or more computer-readable recording media.

[0119] Methods according to embodiments of the present invention may be embodied in the form of program instructions that can be executed by various computer means and stored on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, and the like, alone or in combination. The program instructions stored on the medium may be specially designed and constructed for the embodiments, or may be readily available to those skilled in the art of computer software. Examples of computer-readable storage media include magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specially configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include not only machine code, such as produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform the operations of the embodiments, or vice versa.

[0120] The present invention has been described in detail above with reference to the examples. However, it is obvious that a person having ordinary skill in the art to which the present invention pertains can make various substitutions, additions and modifications within the scope of the technical idea described above, and such modified embodiments also fall within the scope of protection of the present invention as defined by the following claims. [Explanation of symbols]

[0121] 100: System 10: CCTV 20: Information processing server 21: Camera pose estimation unit 22: Mapping target object setting unit 23: Coordinate estimation unit 24: Correction section 25: Message creation and distribution unit 30: Operator terminal

Claims

1. (a) estimating the camera posture by matching road objects in an image captured by a CCTV camera with a precise road map; (b) selecting an object to be mapped within the captured image; and (c) estimating coordinates in the precise road map of a mapping target object selected in the captured image based on the estimated camera pose; The camera posture includes information on the position (X, Y, Z), pan, and tilt of the CCTV camera, The step (a) comprises: (a1) projecting the precise road map onto the photographed image using the estimated camera pose, or back-projecting an object in the photographed image onto the precise road map using the estimated camera pose; (a2) calculating pixel errors of corresponding feature points of the object in the captured image and the object in the precise road map; and (a3) if the calculated pixel error is equal to or greater than a critical value, correcting the estimated camera pose to reduce the calculated pixel error to less than the critical value; The step (c) back-projecting each feature point of the mapping target object onto the precise road map according to the estimated camera pose; and and calculating spatial coordinates of each feature point of the back-projected mapping target object from the precise road map. A method for estimating CCTV camera attitude and object coordinates based on a precise road map.

2. The step (a) comprises: extracting pixel coordinates of feature points of the road surface object from the captured image; extracting spatial coordinates of feature points of the road surface object by matching the road surface object with a precise road map; and and estimating the camera posture using PnP (Perspective-n-Point) based on the relationship between pixel coordinates of the extracted road surface object feature points and spatial coordinates. The method of claim 1, wherein the method is based on a precise road map and includes estimating the CCTV camera pose and object coordinates.

3. If the CCTV camera is a PTZ camera, step (a) includes converting the pan and tilt values ​​provided by the PTZ camera into absolute values ​​based on true north and a vertical direction from the difference between the pan and tilt values ​​of the estimated camera attitude and the pan and tilt values ​​provided by the PTZ camera; and The method further includes reducing the captured image by the inverse of the zoom value provided from the PTZ camera to convert the captured image into an image with a zoom value of '0'. The method of claim 2, wherein the method is for estimating the CCTV camera pose and object coordinates based on a precise road map.

4. In the step (b), The mapping target object may be an area or object automatically recognized and extracted from the captured image according to a predetermined rule based on artificial intelligence (AI), or an area or object manually selected by a user from the captured image. The method of claim 1, wherein the method is based on a precise road map and includes estimating the CCTV camera pose and object coordinates.

Citation Information

Patent Citations

  • Method and device for recognizing road surface identifier and automatic driving vehicle

    CN113887391A

  • Information display device

    JP2007127437A

  • Video processing apparatus, video processing method, and program

    JP2011155477A

  • Traffic light detection device, traffic light detection method and program

    JP2020199994A

  • Position estimation device, vehicle, position estimation method and position estimation program

    JP2021082181A