Visual feature extraction neural network model training method and system

Training a visual feature extraction neural network with distorted top-view images from Street View data addresses the cost and effort of obtaining ground control points, enhancing feature extraction across diverse image domains.

JP2025536926AInactive Publication Date: 2025-11-12NAVER CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025522102
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-18
Filing Date
2023-10-18
Publication Date
2025-11-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Street View images lack absolute coordinate position information, requiring numerous ground markers for obtaining ground control points, which is costly and labor-intensive.

Method used

A method for training a visual feature extraction neural network model using distorted top-view images derived from Street View images, incorporating positional correspondences to enhance feature extraction across different domains.

Benefits of technology

Improves the performance of visual feature extraction neural networks by enabling effective feature extraction between Street View and aerial images, even when they originate from different domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536926000001_ABST
    Figure 2025536926000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a visual feature extraction neural network model training method performed by at least one processor, the visual feature extraction neural network model training method including: receiving a first street view image captured on the ground, transforming the first street view image into a first distorted top-view image, transforming the first street view image into a second distorted top-view image, obtaining a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image, and training a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence, wherein the first distorted top-view image and the second distorted top-view image are different from each other.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method and system for training a visual feature extraction neural network model, and more particularly to a method and system for training a visual feature extraction neural network model using distorted images transformed based on street view images taken on the ground. [Background technology]

[0002] Ground control points are ground control points, i.e., absolute coordinate position information, used to calculate the coordinate transformation formula between the image coordinate system and the map coordinate system. Ground control points can be obtained by measuring the position information of markers installed on the ground using equipment such as a high-precision GPS (global positioning system).

[0003] Meanwhile, with the development of information technology, map information services have been commercialized, and street view images are provided as one type of map information service. For example, a map information service provider may acquire images of an actual space by using a ground vehicle, and then provide the images taken at a specific point on an electronic map as a street view image of the point.

[0004] However, Street View images do not contain absolute coordinate position information, and in order to obtain ground control points for Street View images taken using a mobile vehicle on the ground, a large number of markers must be placed over a wide area, which is problematic in terms of cost and effort. Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure provides a method, a non-transitory computer-readable recording medium having instructions recorded thereon, and an apparatus (system) for solving such problems. [Means for solving the problem]

[0006] The present disclosure can be implemented in numerous ways, including as a method, an apparatus (system), or a computer program stored on a readable storage medium.

[0007] According to one embodiment of the present disclosure, a method for training a visual feature extraction neural network model performed by at least one processor includes the steps of receiving a first street view image taken on the ground; converting the first street view image into a first distorted top-view image; converting the first street view image into a second distorted top-view image; obtaining a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image; and training a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence, wherein the first distorted top-view image and the second distorted top-view image are different from each other.

[0008] A non-transitory computer-readable storage medium having instructions recorded thereon for a computer to perform a method according to one embodiment of the present disclosure is provided.

[0009] According to one embodiment of the present disclosure, an information processing system includes a memory and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory, the at least one program including instructions for receiving a first street view image taken on the ground, transforming the first street view image into a first distorted top-view image, transforming the first street view image into a second distorted top-view image, obtaining a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image, and training a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence, wherein the first distorted top-view image and the second distorted top-view image are different from each other. [Effects of the Invention]

[0010] According to one embodiment of the present disclosure, the performance of a visual feature extraction neural network model can be improved by training the visual feature extraction neural network model using distorted top-view images of street view images.

[0011] According to one embodiment of the present disclosure, by training a visual feature extraction neural network model using distorted top-view images of street view images and aerial images related to the street view images, visual features can be effectively extracted even between images from different domains.

[0012] According to some embodiments of the present disclosure, the performance of a visual feature extraction neural network model can be improved by generating a training dataset using transformation parameters related to virtual road surface data and / or camera viewpoint for street view images, thereby providing high-quality visual feature point matching results even between images from different domains (e.g., street view images and aerial images).

[0013] The effects of the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by a person with ordinary knowledge in the technical field to which the present disclosure pertains (hereinafter referred to as "ordinary engineer") from the description of the claims. [Brief explanation of the drawings]

[0014] Embodiments of the present disclosure will be described, without limitation, with reference to the accompanying drawings described below, in which like reference numerals indicate like elements and in which: [Figure 1] FIG. 1 illustrates an example method for aligning a three-dimensional model with street view data according to one embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating a configuration in which an information processing system according to an embodiment of the present disclosure is communicably connected to a plurality of user terminals. [Figure 3] FIG. 2 is a block diagram illustrating an internal configuration of a user terminal and an information processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates an example of generating training data pairs for a visual feature extraction neural network model using street view images according to one embodiment of the present disclosure. [Figure 5] FIG. 10 illustrates an example of generating a distorted top-view transformation relationship using camera parameters according to an embodiment of the present disclosure. [Figure 6] FIG. 1 illustrates an example of generating training data pairs using street view images and a 3D model for training a visual feature extraction neural network model according to one embodiment of the present disclosure. [Figure 7] FIG. 10 is a diagram illustrating an example of visual feature points matched according to one embodiment of the present disclosure. [Figure 8] 1 is a flowchart illustrating an example method for training a visual feature extraction neural network model according to one embodiment of the present disclosure. [Figure 9] 10 is a flowchart illustrating an example of a visual feature extraction neural network model training method according to another embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the accompanying drawings. However, in the following description, detailed descriptions of well-known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.

[0016] In the accompanying drawings, the same or corresponding components are denoted by the same reference numerals. In addition, in the following description of the embodiments, redundant description of the same or corresponding components may be omitted. However, even if technology related to a component is omitted, it is not intended that such a component is not included in any embodiment.

[0017] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will become clearer with reference to the following embodiments described in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below, and may be realized in various different forms. However, the present embodiments are provided merely to complete the disclosure and to fully convey the scope of the invention to those skilled in the art.

[0018] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification are currently commonly used and general terms that have been selected as much as possible while taking into consideration the functions of the present disclosure. However, these terms may change depending on the intentions or precedents of engineers in the relevant field, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, the meanings thereof will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined based on the meanings of the terms and the overall content of this disclosure, rather than simply by the names of the terms.

[0019] In this specification, the singular expression includes the plural expression unless the context clearly dictates otherwise. Furthermore, the plural expression includes the singular expression unless the context clearly dictates otherwise. When a part in the entire specification includes a certain element, this does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.

[0020] Additionally, the terms "module" and "unit" as used herein refer to software or hardware components, and a "module" or "unit" performs a certain function. However, the term "module" or "unit" is not limited to software or hardware. A "module" or "unit" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, by way of example, a "module" or "unit" may include components such as software components, object-oriented software components, class components, and task components, as well as at least one of processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The components and "modules" or "units" may be combined into fewer components and "modules" or "units," or the functionality provided therein may be further separated into additional components and "modules" or "units."

[0021] According to one embodiment of the present disclosure, a "module" or a "unit" may be implemented with a processor and memory. "Processor" should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, "processor" may refer to an application-specific semiconductor (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. "Processor" may also refer to a combination of processing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or any other such configuration. Furthermore, "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. "Memory" may also refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with a processor if the processor can read information from and / or write information to the memory. Memory that is integrated into a processor is in electronic communication with the processor.

[0022] In the present disclosure, a "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, a system may be configured with one or more server devices. As another example, a system may be configured with one or more cloud devices. As another example, a system may be configured and operated by a server device and a cloud device together.

[0023] In this disclosure, "display" may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by the computing device.

[0024] In the present disclosure, "each of a plurality of A's" or "each of a plurality of A's" may refer to each of all the components included in the plurality of A's, or may refer to each of some of the components included in the plurality of A's.

[0025] In this disclosure, "street view data" may refer to data including not only road view data including images and location information taken on roadways, but also pedestrian view data including images and location information taken on sidewalks. Furthermore, "street view data" may further include not only roadways and sidewalks, but also images and location information taken at any point outdoors (or indoors with a view of the outdoors).

[0026] 1 illustrates an example method for aligning a 3D model 110 and street view data 120 according to one embodiment of the present disclosure. An information processing system may acquire / receive the 3D model 110 and street view data 120 for a particular area.

[0027] The 3D model 110 may include 3D geometric information expressed in absolute coordinate positions and corresponding texture information. Here, the position information included in the 3D model 110 may be information with higher accuracy than the position information included in the street view data 120. Furthermore, the texture information included in the 3D model 110 may be information with lower quality (e.g., lower resolution) than the texture information included in the street view data 120. According to one embodiment, the 3D geometric information expressed in absolute coordinate positions may be generated based on an aerial photograph taken above the specific area.

[0028] The 3D model 110 for a specific area may include a 3D building model 112, a digital elevation model (DEM) 114, a true ortho image 116 for the specific area, a digital surface model (DSM), a road layout, a road DEM, etc. As a specific example, the 3D model 110 for a specific area may be, but is not limited to, a model generated based on a digital surface model (DSM) including geometric information about the ground of the specific area and the corresponding true ortho image 116 for the specific area. In one embodiment, a precise true ortho image 116 for a specific area may be generated based on multiple aerial photographs and absolute coordinate position information and direction information for each aerial photograph.

[0029] The street view data 120 may include a plurality of street view images captured at a plurality of nodes within a specific area and absolute coordinate position information for each of the plurality of street view images. Here, the position information included in the street view data 120 may be less accurate than the position information included in the 3D model 110, and the texture information included in the street view images may be higher quality (e.g., higher resolution) than the texture information included in the 3D model 110. For example, the position information included in the street view data 120 may be position information obtained using a GPS device when capturing street view images at the nodes. Position information obtained using a vehicle's GPS device may have an error of approximately 5 to 10 meters. Furthermore, the street view data may include direction information for each of the plurality of street view images (i.e., image capture direction information).

[0030] The information processing system may perform map matching 130 between the 3D model 110 and the street view data 120. Specifically, the information processing system may perform feature matching between texture information included in the 3D model 110 and a plurality of street view images included in the street view data 120. To perform map matching 130, the information processing system may convert at least some of the plurality of street view images included in the street view data 120 into top view images. As a result of map matching 130, a plurality of map matching points / map matching lines 132 may be extracted.

[0031] A map matching point may represent a correspondence pair between one point in a street view image and one point in the 3D model 110. The types of map matching points may vary depending on the type of 3D model 110 used in map matching 130, the location of the points, and other factors. For example, a map matching point may include at least one of a ground control point (GCP), which is a correspondence pair of points on the ground within a specific area, a building control point (BCP), which is a correspondence pair of points on a building within a specific area, or a structure control point, which is a correspondence pair of points on a structure within a specific area. Map matching points can be extracted from any area of ​​the street view image and the 3D model 110, in addition to the above-mentioned ground, buildings, and structures.

[0032] A map matching line may represent a corresponding pair of a line in a street view image and a line in the 3D model 110. The types of map matching lines may vary depending on the type of 3D model 110 used in map matching 130, the position of the line, and other factors. For example, the map matching line may include at least one of a ground control line (GCL), which is a corresponding pair of lines on the ground within a specific area; a building control line (BCL), which is a corresponding pair of lines on a building within a specific area; a structure control line, which is a corresponding pair of lines on a structure within a specific area; or a lane control line, which is a corresponding pair of lines on a lane within a specific area. Map matching lines can be extracted from any area of ​​the street view image and the 3D model 110, in addition to the above-mentioned ground, buildings, structures, and lanes.

[0033] The information processing system may also perform feature matching 150 between multiple street view images to extract multiple feature point correspondence sets 152. According to one embodiment, for robust feature matching, feature matching 150 between multiple street view images may be performed using at least a portion of the 3D model 110. For example, feature matching 150 between street view images may be performed using 3D building models 112 included in the 3D model 110.

[0034] The information processing system may then estimate (160) absolute coordinate position information and direction information for the plurality of street view images based on at least one of the plurality of map matching points / map matching lines 132 and at least a portion of the plurality of feature point correspondence sets 152. For example, the processor may estimate (160) the absolute coordinate position information and direction information for the plurality of street view images using a bundle adjustment technique. According to one embodiment, the estimated absolute coordinate position information and direction information 162 is information in an absolute coordinate system representing the three-dimensional model 110 and may be six degrees of freedom (DoF) parameters. The absolute coordinate position information and direction information 162 estimated through this process may be data with higher accuracy than the absolute coordinate position information and direction information included in the street view data 120.

[0035] Hereinafter, a method for training a visual feature extraction neural network model used in the process in which the information processing system performs map matching 130 between a three-dimensional model 110 and street view data 120 will be described in detail with reference to FIGS. 4 to 9.

[0036] According to one embodiment of the present disclosure, a training dataset is generated based on street view data 120 captured on the ground and / or aerial images acquired from the three-dimensional model 110, and then a visual feature extraction neural network model is trained. This allows for effective extraction of visual features even between images from different domains (e.g., street view images and aerial images), thereby improving the performance of the visual feature extraction neural network model and providing high-quality visual feature point matching results.

[0037] 2 is a schematic diagram illustrating a configuration in which an information processing system 230 according to an embodiment of the present disclosure is communicatively connected to multiple user terminals 210_1, 210_2, and 210_3. As illustrated, the multiple user terminals 210_1, 210_2, and 210_3 can be connected to the information processing system 230, which can provide a map information service, via a network 220. Here, the multiple user terminals 210_1, 210_2, and 210_3 may include terminals of users who receive the map information service. Furthermore, the multiple user terminals 210_1, 210_2, and 210_3 may be automobiles that capture street view images at nodes. In one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more cloud computing service-based distributed computing devices and / or distributed databases, that can store, provide, and execute computer-executable programs (e.g., downloadable applications) and data related to the provision of the map information service.

[0038] The map information service provided by the information processing system 230 may be provided to users via an application, a web browser, or the like installed on each of the multiple user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 may provide information corresponding to a street view image request, an image-based location recognition request, or the like received from the user terminals 210_1, 210_2, and 210_3 via an application, or may perform corresponding processing.

[0039] A plurality of user terminals 210_1, 210_2, and 210_3 can communicate with the information processing system 230 via a network 220. The network 220 can be configured to enable communication between the plurality of user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 can be configured with, for example, a wired network such as Ethernet, a wired home network (Power Line Communication), a telephone line communication device, and RS-serial communication, a mobile communication network, a wireless network such as WLAN (Wireless LAN), Wi-Fi, Bluetooth, and ZigBee, or a combination thereof. The communication method is not limited, and can include not only a communication method utilizing a communication network that the network 220 can include (for example, a mobile communication network, a wired Internet, a wireless Internet, a broadcast network, a satellite network, etc.), but also short-range wireless communication between the user terminals 210_1, 210_2, and 210_3.

[0040] 2, a mobile terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 are shown as examples of user terminals, but are not limited thereto, and the user terminals 210_1, 210_2, and 210_3 may be any computing devices capable of wired and / or wireless communication and capable of installing and executing an application or a web browser, etc. For example, the user terminals may include AI speakers, smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), tablet PCs, game consoles, wearable devices, internet of things (IoT) devices, virtual reality (VR) devices, augmented reality (AR) devices, set-top boxes, etc. Also, while FIG. 2 shows three user terminals 210_1, 210_2, and 210_3 communicating with the information processing system 230 via the network 220, this is not limited thereto, and a different number of user terminals may be configured to communicate with the information processing system 230 via the network 220.

[0041] According to one embodiment, the information processing system 230 may receive street view images captured on the ground and aerial images related to the street view images from the user terminals 210_1, 210_2, and 210_3. The information processing system 230 may then convert the received street view images into a pair of different distorted top-view images and obtain a positional correspondence between the pair of different distorted top-view images. The information processing system 230 may then train a visual feature extraction neural network model using the pair of different distorted top-view images and the positional correspondence. The information processing system 230 may then perform feature matching using the trained visual feature extraction neural network model. The information processing system 230 may then obtain a plurality of ground control points based on the received street view images and transmit the obtained plurality of ground control points to the user terminals 210_1, 210_2, and 210_3. In addition, the information processing system 230 can transmit various service-related data to the user terminals 210_1, 210_2, and 210_3 based on data created by matching the 3D model with street view data using multiple ground control points.

[0042] 3 is a block diagram illustrating the internal configuration of a user terminal 210 and an information processing system 230 according to an embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of executing an application, a web browser, or the like, and capable of wired / wireless communication, and may include, for example, the mobile phone terminal 210_1, the tablet terminal 210_2, and the PC terminal 210_3 of FIG. 2 . As illustrated, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As illustrated in FIG. 3 , the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data over the network 220 using their respective communication modules 316 and 336. Additionally, the input / output device 320 may be configured to input information and / or data to the user terminal 210 or output information and / or data generated from the user terminal 210 via the input / output interface 318 .

[0043] The memories 312 and 332 may include any non-transitory computer-readable recording medium. According to one embodiment, the memories 312 and 332 may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid state drive (SSD), or a flash memory. As another example, the permanent mass storage device such as a ROM, SSD, flash memory, or disk drive may be a separate permanent storage device distinct from the memory and may be included in the user terminal 210 or the information processing system 230. The memories 312 and 332 may also store an operating system and at least one program code (e.g., code for an application installed and run on the user terminal 210).

[0044] These software components may be loaded from a computer-readable recording medium separate from the memories 312, 332. Such separate computer-readable recording medium may include a recording medium directly connectable to the user terminal 210 and the information processing system 230, such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card. As another example, the software components may be loaded into the memories 312, 332 via a communication module rather than a computer-readable recording medium. For example, at least one program may be loaded into the memories 312, 332 based on a computer program to be installed by a file provided via the network 220 by a developer or a file distribution system that distributes application installation files.

[0045] The processors 314, 334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to the processors 314, 334 by the memories 312, 332 or the communications modules 316, 336. For example, the processors 314, 334 may be configured to execute received instructions according to program code stored in a storage device, such as the memories 312, 332.

[0046] The communication modules 316, 336 may provide a configuration or function for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and may provide a configuration or function for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (e.g., a separate cloud system). For example, a request or data (e.g., data related to a request for street view images and aerial images taken on the ground) generated by the processor 314 of the user terminal 210 in accordance with program code stored in a storage device such as the memory 312 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, a control signal or command provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 316 of the user terminal 210 via the communication module 336 and the network 220. For example, the user terminal 210 may receive data related to street view images and aerial images of a specific area from the information processing system 230.

[0047] The input / output interface 318 may be a means for interfacing with the input / output device 320. As an example, the input device may include a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interfacing with a device that integrates input and output configurations or functions, such as a touchscreen. For example, when the processor 314 of the user terminal 210 processes instructions of a computer program loaded into the memory 312, a service screen configured using information and / or data provided by the information processing system 230 or another user terminal may be displayed on the display via the input / output interface 318. Although FIG. 3 illustrates the input / output device 320 as not being included in the user terminal 210, this is not limiting and the input / output device 320 may be configured as a single device together with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be a means for interfacing with an input or output device (not shown) that is connected to the information processing system 230 or that may be included in the information processing system 230. In Fig. 3, the input / output interfaces 318, 338 are shown as elements configured separately from the processors 314, 334, but are not limited to this, and the input / output interfaces 318, 338 may be configured to be included in the processors 314, 334.

[0048] The user terminal 210 and the information processing system 230 may include more components than those shown in FIG. 3 . However, it is not necessary to explicitly illustrate most of the components of the prior art. According to one embodiment, the user terminal 210 may be implemented to include at least a portion of the input / output device 320 described above. The user terminal 210 may also include other components such as a transceiver, a Global Positioning System (GPS) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, it may include components typically included in a smartphone, such as an acceleration sensor, a gyro sensor, an image sensor, a proximity sensor, a touch sensor, an illuminance sensor, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration. According to one embodiment, the processor 314 of the user terminal 210 may be configured to run an application that provides a map information service. In this case, code related to the application and / or program may be loaded into the memory 312 of the user terminal 210.

[0049] While a program for an application that provides a map information service is running, the processor 314 may receive text, images, pictures, voice, and / or actions entered or selected via an input device such as a touch screen, keyboard, camera including an audio sensor and / or image sensor, or microphone connected to the input / output interface 318, and may store the received text, images, pictures, voice, and / or actions in the memory 312 or provide them to the information processing system 230 via the communication module 316 and the network 220. For example, the processor 314 may receive a user input requesting street view images and aerial images for a specific area and provide them to the information processing system 230 via the communication module 316 and the network 220.

[0050] The processor 314 of the user terminal 210 may be configured to manage, process, and / or store information and / or data received from the input / output device 320, other user terminals, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 may be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the user terminal 210 may transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 may display the received information and / or data on the screen of the user terminal.

[0051] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from multiple user terminals 210 and / or multiple external systems. The information and / or data processed by the processor 334 may be provided to the user terminals 210 via the communication module 336 and the network 220.

[0052] FIG. 4 illustrates an example of generating training data pairs 440 for a visual feature extraction neural network model using street view images 410 according to an embodiment of the present disclosure. As illustrated, an information processing system can generate training data pairs 440 for a visual feature extraction neural network model by converting street view images 410 into different distorted top view images. Here, the street view images 410 may be images captured on the ground using a vehicle equipped with at least one camera. For example, the street view images 410 may be 360-degree panoramic images captured on a road on the ground. Information such as location information and pose information about the street view images 410 is not required to generate the training data pairs 440.

[0053] Training data pair 440 can be generated using a first distorted top-view image and a second distorted top-view image that have been transformed based on street view image 410. In one example, training data pair 440 can include the first distorted top-view image, the second distorted top-view image, and positional correspondences between pixels in the first distorted top-view image and pixels in the second distorted top-view image.

[0054] According to one embodiment, the information processing system may acquire (422) first virtual road surface data and (432) second virtual road surface data. Here, the first virtual road surface data and the second virtual road surface data may be different from each other. For example, the information processing system may acquire the first virtual road surface data and the second virtual road surface data by assuming that the road surface is curved rather than flat and selecting two of a large number of pre-generated virtual road surface data. One of the parameters that causes distortion when converting street view images captured by a real vehicle into virtual top-view images is the road profile (i.e., road surface data) rather than a flat surface. Specifically, because the road surface is not perfectly flat when a real vehicle captures street view images, distortion occurs when converting the street view image 410 into a virtual top-view image. By generating a distorted top-view image that reflects distortion based on such parameters and training a visual feature extraction neural network model using training data pairs based on the distorted top-view image, matching performance can be efficiently improved.

[0055] According to one embodiment, the information processing system may acquire (424) first camera parameters and (434) second camera parameters. Here, the camera parameters may be image transformation parameters related to the viewpoint / attitude of the camera. In one embodiment, the camera parameters may include, but are not limited to, a height value, a pitch value, and a roll value between the road and the camera. The first camera parameter and the second camera parameter may be different from each other. That is, the first camera parameter and the second camera parameter may be different from each other by setting at least one value of the height value, pitch value, and roll value between the road and the camera to be different. For example, the information processing system may acquire the first camera parameter and the second camera parameter by selecting two from a number of pre-generated camera parameters. One of the parameters that causes distortion when converting a street view image captured by an actual vehicle into a virtual top-view image is the camera parameter. Here, the camera parameter refers to the camera mounting height, pitch angle, and roll angle relative to the road. By generating distorted top-view images that reflect the distortion based on such parameters and training data pairs based on these images to train a visual feature extraction neural network model, matching performance can be efficiently improved.

[0056] The information processing system may then convert the street view image 410 into a top-view image using the first virtual road surface data and the first camera parameters to generate a first distorted top-view image (426). Similarly, the information processing system may convert the street view image 410 into a top-view image using the second virtual road surface data and the second camera parameters to generate a second distorted top-view image (436).

[0057] The information processing system can then obtain a positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image, where the positional correspondence can be determined based on the first virtual road surface data, the first camera parameters, the second virtual road surface data, and the second camera parameters. Specifically, the information processing system can obtain a positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image using transformation relationship data based on the first virtual road surface data and the first camera parameters and transformation relationship data based on the second virtual road surface data and the second camera parameters.

[0058] The information processing system can then generate training data pairs 440 using the first distorted top-view image, the second distorted top-view image, and the positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image. For example, one pixel in the first distorted top-view image and one pixel in the second distorted top-view image that has the same positional correspondence as the first pixel can be used as training data pair 440. In another example, the information processing system can use only pixels in the first distorted top-view image and / or the second distorted top-view image that have a confidence score (or importance score) equal to or greater than a predetermined threshold as training data pair 440. In this case, the visual feature extraction neural network model can be a model that outputs an N-dimensional visual feature descriptor and a confidence score for each pixel of the input image. The generated training data pairs 440 can be used to train the visual feature extraction neural network model.

[0059] In one embodiment, the visual feature extraction neural network model can be trained to make visual feature descriptors of corresponding pixels in the first distorted top-view image and the second distorted top-view image similar. For example, the visual feature extraction neural network model can be a model in which a network training loss is configured to make visual feature descriptors of each pixel similar using the correspondence between the first distorted top-view image and the second distorted top-view image.

[0060] 4 illustrates, but is not limited to, a process in which a pair of distorted top-view images is generated based on one street view image 410, and training data pair 440 is generated based on the distorted top-view images. For example, the information processing system may generate various training data pairs by changing the settings of virtual road surface data and / or camera parameters and applying them to street view image 410. Furthermore, the information processing system may generate additional training data pairs by applying various virtual road surface data and / or camera parameters to multiple street view images. The method of generating training data pair 440 illustrated in FIG. 4 can be used when there is no aerial image corresponding to street view image 410.

[0061] 4 illustrates applying both virtual road surface data and camera parameters to a street view image to generate a distorted top-view image, but this is not limiting. For example, the distorted top-view image can be generated by applying only virtual road surface data to a street view image. In another example, the distorted top-view image can be generated by applying only camera parameters to a street view image.

[0062] 5 illustrates an example of generating a distorted top-view transformation relationship 550 using virtual road surface data 510 and camera parameters 520, 530, and 540 according to one embodiment of the present disclosure. As illustrated, the distorted top-view transformation relationship 550 can be generated by applying the virtual road surface data 510 and the camera parameters 520, 530, and 540. A distorted top-view image can be generated by converting a street-view image into a virtual top-view image using the top-view transformation relationship 550. In this manner, using the top-view transformation relationship 550 can mimic the distortion that occurs when converting a street-view image captured by an actual vehicle into a virtual top-view image.

[0063] In one embodiment, the virtual road surface data 510 may be generated by assuming that the shape of the road included in the street view image is a curved surface rather than a flat surface. For example, the virtual road surface data 510 may be data that represents horizontal curves, vertical curves, etc. in the form of spline curves.

[0064] In one embodiment, the camera parameters 520, 530, and 540 may be image transformation parameters related to the camera's viewpoint. Since a street view image is an image captured using a vehicle equipped with a single camera, a height value 530, a pitch value 520, and a roll value 540 between the road and the camera can be used as transformation parameters. Here, the pitch value 520 may represent a value (horizontal axis rotation value) of the camera rotated up / down relative to a state in which the road and the camera are parallel. The height value 530 may represent a value of the height between the road and the camera's viewpoint. Furthermore, the roll value 540 may represent a value (vertical axis rotation value) of the camera rotated left / right relative to a state in which the road and the camera are parallel. The camera parameters may be generated / acquired using a combination of the height value 530, pitch value 520, and roll value 540 between the road and the camera.

[0065] By using the above-described configuration, it is possible to generate a variety of high-quality effective training data sets compared to homography transformation, which has a high degree of freedom and includes transformation data that is not effective for training a visual feature extraction neural network model.

[0066] 6 illustrates an example of generating training data pairs 670 for training a visual feature extraction neural network model using street view imagery 610 and a 3D model 650 according to an embodiment of the present disclosure. Here, the 3D model 650 may be a model of a specific area generated based on aerial imagery. In one example, the 3D model 650 may include three-dimensional absolute coordinate position information (e.g., a DEM) of the specific area and a true ortho image 652.

[0067] Training data pairs 670 can be generated using distorted top-view images and true orthoimages 652 that have been transformed based on street view images 610. In one embodiment, training data pairs 670 can include distorted top-view images, true orthoimages 652, and positional correspondences between pixels in the distorted top-view images and pixels in the true orthoimages 652.

[0068] According to one embodiment, the information processing system may acquire (620) virtual road surface data. For example, the information processing system may acquire the virtual road surface data by assuming that the road surface is curved rather than flat and selecting one of a plurality of pre-generated virtual road surface data. The information processing system may then acquire (630) camera parameters. For example, the information processing system may acquire the camera parameters by selecting one of a number of pre-generated camera parameters. The information processing system may then convert (640) the street view image 610 into a top-view image using the virtual road surface data and the camera parameters.

[0069] In one embodiment, the information processing system can retrieve a true orthoimage 652 corresponding to the street view image 610 from the 3D model 650. For example, high-precision / high-accuracy 6DoF (Degree of Freedom) pose information associated with the street view image 610 can be used to retrieve an aerial image (e.g., the true orthoimage 652) of an area associated with the street view image 610 from the 3D model 650. For example, the high-precision / high-accuracy 6DoF pose information associated with the street view image 610 can be obtained by post-processing low-precision / low-accuracy 6DoF pose information acquired when capturing the street view image 610 with information obtained from a high-quality sensor (e.g., RTK-GPS, LiDAR, IMU, wheel odometry, etc.).

[0070] The information processing system can then obtain positional correspondences between pixels in the distorted top-view image and pixels in the true orthoimage 652, where the positional correspondences can be determined by calculating 660 optical flow between the distorted top-view image and the true orthoimage.

[0071] The information processing system can then generate training data pairs 670 using the distorted top-view image, the true orthoimage, and the positional correspondence between pixels in the distorted top-view image and pixels in the true orthoimage 652. For example, one pixel in the distorted top-view image and one pixel in the true orthoimage 652 that has the same positional correspondence as the pixel in the distorted top-view image can be used as the training data pair 670. In another example, the information processing system can use only pixels in the distorted top-view image and / or the true orthoimage 652 that have a confidence score (or importance score) equal to or greater than a predetermined threshold as the training data pair 670. In this case, the visual feature extraction neural network model can be a model that outputs an N-dimensional visual feature descriptor and a confidence score for each pixel of the input image. The generated training data pairs 670 can be used to train the visual feature extraction neural network model.

[0072] In one embodiment, the visual feature extraction neural network model can be trained to make the visual feature descriptors of corresponding pixels in the distorted top-view image and the true orthoimage 652 similar. For example, the visual feature extraction neural network model can be a model in which a network training loss is configured to make the visual feature descriptors of each pixel similar using the correspondence between the distorted top-view image and the true orthoimage.

[0073] 6 illustrates, but is not limited to, a process of generating a distorted top-view image based on one street view image 610 and generating training data pair 670 based on the corresponding true orthoimage 652. For example, the information processing system can generate various training data pairs by changing the settings of the virtual road surface data and / or camera parameters in various ways and applying the changes to the street view image 610. Furthermore, the information processing system can generate additional training data pairs by applying various virtual road surface data and / or camera parameters to multiple street view images. The method of generating training data pair 670 illustrated in FIG. 6 can be used when there is an aerial image (e.g., image 652) corresponding to the street view image 610.

[0074] In one embodiment, when a plurality of Street View images do not have corresponding aerial images, the method for generating training data pairs shown in FIG. 4 is used to generate a large amount of training data. When a plurality of Street View images have corresponding aerial images, the method for generating training data pairs shown in FIG. 6 is used to generate a large amount of training data.

[0075] 7 is a diagram illustrating an example of visual feature points matched according to an embodiment of the present disclosure. Feature matching can be performed by converting a street view image into a virtual top-view image and then using the converted virtual top-view image and its corresponding true orthoimage. In this case, feature matching can be performed using a visual feature extraction neural network model trained using training data generated by the above-described method.

[0076] A first graph 710 shows an example of the results of feature matching between a virtual top-view image and its corresponding true orthoimage using a conventional visual feature extractor. A second graph 720 shows an example of the results of feature matching between the same virtual top-view image and its corresponding true orthoimage using a visual feature extraction neural network model trained by the training method of the present disclosure. As shown, it can be seen that the visual feature extraction neural network model of the present disclosure, trained using virtual road surface data and / or camera parameters, has a reduced mismatch rate compared to conventional visual feature extractors, and can provide more accurate feature matching results for images from different domains (virtual top-view images converted from street view images and aerial images (e.g., true orthoimages)).

[0077] 8 is a flowchart illustrating an example method 800 for training a visual feature extraction neural network model according to one embodiment of the present disclosure. Method 800 may begin with a processor (e.g., at least one processor of an information processing system) receiving street view images captured on the ground (S810). The processor may then convert the first street view image into a first distorted top-view image (S820). The processor may also convert the first street view image into a second distorted top-view image (S830). Here, the first distorted top-view image and the second distorted top-view image may be different from each other.

[0078] In one embodiment, the processor can convert the street view image into a distorted top-view image using the virtual road surface data. For example, the processor can obtain first virtual road surface data and convert the first street view image into a first distorted top-view image using the first virtual road surface data. The processor can also obtain second virtual road surface data and convert the first street view image into a second distorted top-view image using the second virtual road surface data. Here, the first virtual road surface data and the second virtual road surface data can be different from each other.

[0079] In one embodiment, the processor may convert the street view image into a distorted top-view image using the camera parameters. For example, the processor may acquire first camera parameters and convert the first street view image into a first distorted top-view image using the first camera parameters. The processor may also acquire second camera parameters and convert the first street view image into a second distorted top-view image using the second camera parameters. Here, the first camera parameters and the second camera parameters may be different from each other. Here, the first camera parameters and the second camera parameters may include a height value, a pitch value, and a roll value between the road and the camera, respectively.

[0080] In one embodiment, the processor can convert the street view image into a distorted top-view image using the virtual road surface data and the camera parameters. For example, the processor can acquire first virtual road surface data and acquire first camera parameters. The processor can then convert the first street view image into a first distorted top-view image using the first virtual road surface data and the first camera parameters. Similarly, the processor can acquire second virtual road surface data and acquire second camera parameters. The processor can then convert the first street view image into a second distorted top-view image using the second virtual road surface data and the second camera parameters. Here, the first virtual road surface data and the second virtual road surface data can be different from each other. Additionally or alternatively, the first camera parameters and the second camera parameters can be different from each other.

[0081] The processor may obtain a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image (S840), where the first positional correspondence may be determined based on the first virtual road surface data, the first camera parameters, the second virtual road surface data, and the second camera parameters.

[0082] The processor may then train a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence (S850). Here, the visual feature extraction neural network model may be trained so that visual feature descriptors between corresponding pixels in the first distorted top-view image and the second distorted top-view image are similar. Alternatively, the visual feature extraction neural network model may be trained so that visual feature descriptors of corresponding pixels in the first distorted top-view image and the second distorted top-view image that have confidence scores above a predetermined threshold are similar.

[0083] The learning method 800 of Figure 8 can be used when there is no aerial imagery corresponding to the street view imagery. The flowchart of Figure 8 and the above description are merely examples, and the scope of the present disclosure is not limited thereto. For example, at least one step may be added / modified / deleted, or the order of the steps may be changed.

[0084] 9 is a flowchart illustrating an example of a visual feature extraction neural network model training method 900 according to another embodiment of the present disclosure. The method 900 may begin by a processor (e.g., at least one processor of an information processing system) receiving street view images captured on the ground (S910). The processor may then convert the street view images into distorted top-view images (e.g., a third distorted top-view image) (S920).

[0085] In one embodiment, the processor can obtain virtual road surface data and camera parameters. The processor can then use the virtual road surface data and the camera parameters to convert the street view image into a distorted top-view image. In another example, the processor can use the virtual road surface data to convert the street view image into a distorted top-view image. In another example, the processor can use the camera parameters to convert the street view image into a distorted top-view image.

[0086] The processor may then receive an aerial image associated with the street view image (S930), where the aerial image may be a true-ortho image. The processor may then obtain a positional correspondence between pixels in the distorted top-view image and the aerial image (S940), where the positional correspondence may be determined based on optical flow calculated based on the distorted top-view image and the aerial image.

[0087] The processor may train a visual feature extraction neural network model using the distorted top-view image, the aerial image, and the positional correspondences (S950), where the visual feature extraction neural network model may be trained to similarly represent visual feature descriptors between corresponding pixels in the distorted top-view image and the aerial image.

[0088] The training method 900 of FIG. 9 can be used when there are aerial images (e.g., true orthoimages) corresponding to street view images. The flowchart of FIG. 9 and the above description are merely examples, and the scope of the present disclosure is not limited thereto. For example, at least one step may be added / modified / deleted, or the order of the steps may be changed.

[0089] The above-described method can be provided as a computer program stored on a computer-readable recording medium for execution by a computer. The medium may permanently store a computer-executable program or may temporarily store the program for execution or download. The medium may be various recording or storage means combined with one or more pieces of hardware, and is not limited to media directly connected to a computer system but may also be distributed over a network. Examples of media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Other examples of media include recording or storage media managed by app stores that distribute applications, or by sites or servers that provide or distribute various other software.

[0090] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, the techniques may be implemented in hardware, firmware, software, or a combination thereof. Those of ordinary skill in the art will understand that the various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may also be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design requirements for the overall system. Those of ordinary skill in the art may implement the described functionality in various ways for each particular application, but such implementation should not be interpreted as causing a departure from the scope of the present disclosure.

[0091] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or combinations thereof.

[0092] Accordingly, the various illustrative logic blocks, modules, and circuits described in connection with this disclosure may be implemented or performed with a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other configuration.

[0093] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processors to perform certain aspects of the functions described in this disclosure.

[0094] While the above-described embodiments have been described as utilizing aspects of the presently disclosed subject matter on one or more stand-alone computer systems, the present disclosure is not limited thereto and may be implemented in connection with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the subject matter in this disclosure may be implemented on multiple processing chips or devices, and storage may be affected similarly across multiple devices. These devices may include PCs, network servers, and handheld devices.

[0095] Although the present disclosure has been described herein with reference to several embodiments, various modifications and changes can be made without departing from the scope of the present disclosure, which would be understood by those of ordinary skill in the art to which the presently disclosed invention pertains, and these modifications and changes should also be understood to fall within the scope of the claims appended hereto.

Claims

1. 1. A visual feature extraction neural network model training method performed by at least one processor, comprising: receiving a first street view image taken on the ground; transforming the first street view image into a first distorted top-view image; transforming the first street view image into a second distorted top-view image; obtaining a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image; training a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence; the first distorted top-view image and the second distorted top-view image are different from each other.

2. Transforming the first street view image into a first distorted top-view image includes: acquiring first virtual road surface data; and transforming the first street view image into the first distorted top-view image using the first virtual road surface data; Transforming the first street view image into a second distorted top-view image includes: acquiring second virtual road surface data; and transforming the first street-view image into the second distorted top-view image using the second virtual road surface data; 2. The visual feature extraction neural network model training method according to claim 1, wherein the first virtual road surface data and the second virtual road surface data are different from each other.

3. Transforming the first street view image into a first distorted top-view image includes: obtaining first camera parameters; and transforming the first street view image into the first distorted top-view image using the first camera parameters; Transforming the first street view image into a second distorted top-view image includes: obtaining second camera parameters; and transforming the first street view image into the second distorted top-view image using the second camera parameters; The visual feature extraction neural network model training method of claim 1 , wherein the first camera parameters and the second camera parameters are different from each other.

4. 4. The visual feature extraction neural network model training method according to claim 3, wherein the first camera parameter and the second camera parameter respectively include a height value, a pitch value, and a roll value between a road and a camera.

5. Transforming the first street view image into a first distorted top-view image includes: acquiring first virtual road surface data; obtaining first camera parameters; and transforming the first street view image into the first distorted top-view image using the first virtual road surface data and the first camera parameters; Transforming the first street view image into a second distorted top-view image includes: acquiring second virtual road surface data; obtaining second camera parameters; and transforming the first street view image into the second distorted top-view image using the second virtual road surface data and the second camera parameters; the first virtual road surface data and the second virtual road surface data are different from each other; The visual feature extraction neural network model training method of claim 1 , wherein the first camera parameters and the second camera parameters are different from each other.

6. 6. The visual feature extraction neural network model training method according to claim 5, wherein the first positional correspondence relationship is determined based on the first virtual road surface data, the first camera parameters, the second virtual road surface data, and the second camera parameters.

7. 2. The method of training a visual feature extraction neural network model according to claim 1, wherein the visual feature extraction neural network model is trained such that visual feature descriptors of corresponding pixels in the first distorted top-view image and the second distorted top-view image are similar.

8. 2. The method of training a visual feature extraction neural network model according to claim 1, wherein the visual feature extraction neural network model is trained such that visual feature descriptors between corresponding pixels in the first distorted top-view image and the second distorted top-view image that have confidence scores equal to or greater than a predetermined threshold are similar.

9. receiving a second street view image taken on the ground; transforming the second street view image into a third distorted top-view image; receiving an aerial image associated with the second street view image; obtaining a second positional correspondence between pixels in the third distorted top-view image and the aerial image; 2. The method of claim 1, further comprising: training the visual feature extraction neural network model using the third distorted top-view image, the aerial image, and the second positional correspondence.

10. The visual feature extraction neural network model training method according to claim 9 , wherein the aerial images are true orthoimages.

11. 10. The method of claim 9, wherein the second positional correspondence is determined based on an optical flow between the third distorted top-view image and the aerial image.

12. Transforming the second street view image into a third distorted top-view image includes: acquiring third virtual road surface data; and converting the second street view image into the third distorted top-view image using the third virtual road surface data.

13. Transforming the second street view image into a third distorted top-view image includes: obtaining a third camera parameter; and transforming the second street view image into the third distorted top-view image using the third camera parameters.

14. Transforming the second street view image into a third distorted top-view image includes: acquiring third virtual road surface data; obtaining a third camera parameter; and transforming the second street view image into the third distorted top-view image using the third virtual road surface data and the third camera parameters.

15. A non-transitory computer-readable storage medium having recorded thereon instructions for performing the method of claim 1 on a computer.

16. An information processing system, Memory and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory; receiving a first street view image captured on the ground; transforming the first street view image into a first distorted top-view image; transforming the first street view image into a second distorted top-view image; obtaining a first positional correspondence between pixels in the first distorted top-view image and pixels in the second distorted top-view image; instructions for training a visual feature extraction neural network model using the first distorted top-view image, the second distorted top-view image, and the first positional correspondence; The first distorted top-view image and the second distorted top-view image are different from each other.

17. 1. A visual feature extraction neural network model training method performed by at least one processor, comprising: receiving street view images taken on the ground; converting the street view image into a distorted top-view image; receiving aerial imagery related to the street view imagery; obtaining a positional correspondence between pixels in the distorted top-view image and the aerial image; and training a visual feature extraction neural network model using the distorted top-view image, the aerial image, and the positional correspondences.

18. 18. The visual feature extraction neural network model training method according to claim 17, wherein the positional correspondence is determined based on an optical flow calculated based on the distorted top-view image and the aerial image.

19. converting the street view image into a distorted top-view image, acquiring virtual road surface data; obtaining camera parameters; and converting the street view image into the distorted top-view image using the virtual road surface data and the camera parameters.

20. 18. The method for training a visual feature extraction neural network model according to claim 17, wherein the visual feature extraction neural network model is trained so that visual feature descriptors between corresponding pixels of the distorted top-view image and the aerial image are similar.

Citation Information

Patent Citations

  • Pavement crack detection method and system based on deep learning network

    CN115082450A

  • Method and system for detecting changes in road-layout information

    US20210333124A1