Image processing methods, apparatus, devices, computer programs and storage media
By extracting features and performing semantic recognition on road images, the area to be reconstructed and traffic element images are identified, and 3D reconstruction is performed. This solves the problem of low efficiency in generating real-scene images in existing technologies and achieves more efficient real-scene image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies require significant time and manpower to generate real-world images, resulting in low generation efficiency.
By performing feature extraction and semantic recognition on the road images to be processed, the area to be reconstructed and traffic element images are identified, and three-dimensional reconstruction is performed to generate a three-dimensional real-scene image.
It simplifies the process of generating real-world images, improves processing speed and efficiency, and reduces the workload of image processing.
Smart Images

Figure CN114429528B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to artificial intelligence and vehicle-mounted technology, and particularly relates to an image processing method and device, equipment, a computer program and a storage medium. BACKGROUND
[0002] At present, a real scene enlarged image is usually used in navigation to provide a more intuitive navigation experience for a user. The real scene image is an image of a specific intersection shot from a relevant perspective, and is generated by manually drawing or modeling a relevant scene according to the image and selecting a specific angle for rendering. The current real scene image needs to be drawn or modeled based on an intersection image obtained by surveying, which requires a high time cost and labor cost, thereby reducing the efficiency of generating the real scene image. SUMMARY
[0003] The embodiments of the present application provide an image processing method, device, equipment, computer program and storage medium, which can improve the efficiency of generating a real scene image.
[0004] The technical scheme of the embodiments of the present application is implemented as follows:
[0005] The embodiments of the present application provide an image processing method, comprising:
[0006] The feature extraction and semantic recognition are performed on the to-be-processed road image to obtain a to-be-reconstructed region in the to-be-processed road image and a traffic element image in the to-be-reconstructed region;
[0007] The to-be-reconstructed region is three-dimensionally reconstructed to obtain a three-dimensional scene point cloud;
[0008] The traffic element image is mapped into the three-dimensional scene point cloud according to the type of the traffic element image to generate a three-dimensional real scene image.
[0009] The embodiments of the present application provide an image processing device, comprising:
[0010] The semantic recognition module is configured to perform the feature extraction and semantic recognition on the to-be-processed road image to obtain a to-be-reconstructed region in the to-be-processed road image and a traffic element image in the to-be-reconstructed region;
[0011] The three-dimensional reconstruction module is configured to perform three-dimensional reconstruction on the to-be-reconstructed region to obtain a three-dimensional scene point cloud;
[0012] The generation module is configured to map the traffic element image into the three-dimensional scene point cloud according to the type of the traffic element image to generate a three-dimensional real scene image.
[0013] In the device, the semantic recognition module is further configured to perform feature extraction on the to-be-processed road image to obtain an initial feature point set, perform semantic recognition and segmentation based on the initial feature point set to obtain a region in the to-be-processed road image that is semantically represented as road information, and obtain the traffic element image, the traffic element image being an image region that is semantically represented as traffic indication information; and the region that is semantically represented as road information is taken as the to-be-reconstructed region.
[0014] In the device, the three-dimensional reconstruction module is further configured to extract a reconstruction feature point set corresponding to the to-be-reconstructed region from the to-be-processed road image, and perform three-dimensional reconstruction by using the reconstruction feature point set to obtain the three-dimensional scene point cloud.
[0015] In the device, the three-dimensional reconstruction module is further configured to add a mask to a region outside the to-be-reconstructed region in the to-be-processed road image to obtain a to-be-reconstructed image, and perform feature extraction on the to-be-reconstructed image to obtain the reconstruction feature point set.
[0016] In the device, the image processing apparatus further includes an acquisition module and a fusion module, the acquisition module is configured to acquire a target image of a current road scene and acquire current positioning information corresponding to the current road scene before performing feature extraction and semantic recognition on the to-be-processed road image, and acquire at least one crowd-sourced image according to the current positioning information; and the fusion module is configured to perform image fusion on the at least one crowd-sourced image and the target image to obtain the to-be-processed road image corresponding to the current road scene, the at least one crowd-sourced image being an image collected and uploaded by at least one terminal for the current road scene.
[0017] In the device, the generation module is configured to acquire imaging parameters corresponding to the three-dimensional scene point cloud, calculate an imaging track, calculate a position mapping relationship between the traffic element image and the three-dimensional scene point cloud based on the imaging track, calculate the imaging parameters through a three-dimensional reconstruction process, determine a projection region corresponding to the traffic element image in the three-dimensional scene point cloud according to the position mapping relationship, and map the traffic element image into the projection region according to a type of the traffic element image to generate the three-dimensional real scene map.
[0018] In the device, the type of the traffic element image includes a general type, and the generation module is further configured to, in a case where the traffic element image is of the general type, acquire a vectorized image corresponding to the traffic element image, replace the projection region with the vectorized image, update the three-dimensional scene point cloud, and render the updated three-dimensional scene point cloud based on a preset angle to obtain the three-dimensional real scene map.
[0019] In the apparatus, the generation module is further configured to acquire a target vector template corresponding to the traffic element image from a preset vector library, as the vectorized image; the preset vector library comprises at least one vector template corresponding to at least one traffic element image of a general type; or replace the traffic element image with a grid, and perform vectorization drawing based on the grid to obtain the vectorized image.
[0020] In the apparatus, the generation module is further configured to, in a case where the traffic element image is of a general type, extract an image texture from the traffic element image, add the image texture to the projection area, update the three-dimensional scene point cloud to obtain an updated three-dimensional scene point cloud, and render the updated three-dimensional scene point cloud based on a preset angle to obtain the three-dimensional real scene image.
[0021] In the apparatus, the type of the traffic element image comprises a specific type, and the generation module is further configured to, in a case where the traffic element image is of the specific type, render the three-dimensional scene point cloud or the updated three-dimensional scene point cloud based on a preset angle to generate an initial three-dimensional real scene image, and update the initial three-dimensional real scene image using the traffic element image according to position information of the projection area in the initial three-dimensional real scene image to obtain the three-dimensional real scene image.
[0022] In the apparatus, the traffic element image comprises at least one road line detection point, and the generation module is further configured to perform three-dimensional projection and topological connection on the at least one road line detection point to generate a road line image in the three-dimensional real scene image.
[0023] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:
[0024] A memory is configured to store executable instructions.
[0025] A processor is configured to execute the executable instructions stored in the memory to implement the image processing method provided in the embodiments of the present application.
[0026] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores executable instructions, which are used to cause a processor to execute the image processing method provided in the embodiments of the present application.
[0027] A computer program product is provided in an embodiment of the present application, and the computer program product comprises a computer program or instructions, which are executed by a processor to implement the image processing method provided in the embodiments of the present application.
[0028] The embodiments of the present application have the following beneficial effects:
[0029] By performing semantic recognition on the to-be-processed road image, the to-be-reconstructed region in the to-be-processed road image can be obtained, and then real scene map generation is performed based on the to-be-reconstructed region, thereby reducing the workload of image processing. Moreover, according to the type of the traffic element image, the traffic element image is mapped to the three-dimensional scene point cloud obtained by three-dimensional reconstruction of the to-be-reconstructed region, compared with the related art of drawing and modeling each element in the road image, the generation process of the real scene map is simplified, the processing speed is improved, and thus the efficiency of generating the real scene map is improved. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 a is a navigation effect schematic diagram in the form of a vector diagram provided by an embodiment of the present application;
[0031] Figure 1 b is a navigation effect schematic diagram in the form of a real scene map provided by an embodiment of the present application;
[0032] Figure 2 is an optional system structure schematic diagram of an image processing system provided by an embodiment of the present application;
[0033] Figure 3 is an optional structure schematic diagram of an image processing apparatus provided by an embodiment of the present application;
[0034] Figure 4 is an optional flow schematic diagram of an image processing method provided by an embodiment of the present application;
[0035] Figure 5 is an optional flow schematic diagram of an image processing method provided by an embodiment of the present application;
[0036] Figure 6 is an optional flow schematic diagram of an image processing method provided by an embodiment of the present application;
[0037] Figure 7 is an optional effect schematic diagram of extracting an initial feature point set from a to-be-processed road image provided by an embodiment of the present application;
[0038] Figure 8 is an optional flow schematic diagram of an image processing method provided by an embodiment of the present application;
[0039] Figure 9 is an optional effect schematic diagram of adding a mask to a to-be-processed road image provided by an embodiment of the present application;
[0040] Figure 10 is an optional flow schematic diagram of an image processing method provided by an embodiment of the present application;
[0041] Figure 11is an optional flow diagram of an image processing method provided by an embodiment of the present application;
[0042] Figure 12 is an optional flow diagram of an image processing method provided by an embodiment of the present application;
[0043] Figure 13 is an optional effect diagram of a traffic element image obtained through semantic recognition provided by an embodiment of the present application;
[0044] Figure 14a is an optional effect diagram of a vector template map provided by an embodiment of the present application;
[0045] Figure 14b is an optional effect diagram of a vector template map provided by an embodiment of the present application;
[0046] Figure 14c is an optional effect diagram of a vector template map provided by an embodiment of the present application;
[0047] Figure 15 is an optional effect diagram of a three-dimensional scene point cloud provided by an embodiment of the present application;
[0048] Figure 16 is an optional effect diagram of an initial three-dimensional real scene map generated through vector map replacement provided by an embodiment of the present application;
[0049] Figure 17 is an optional flow diagram of an image processing method provided by an embodiment of the present application;
[0050] Figure 18 is an optional effect diagram of a three-dimensional real scene map generated through an affine transformation method provided by an embodiment of the present application;
[0051] Figure 19 is an optional flow diagram of an image processing method provided by an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0053] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0054] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0055] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0057] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0058] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0059] 1) Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science used to capture the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. Furthermore, AI is used to study the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. Moreover, AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning (ML) / deep learning.
[0060] 2) Computer Vision (CV) Technology: Computer vision is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing, recording, and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0061] 3) Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0062] 4) Autonomous driving technology typically includes high-precision maps, environmental perception, behavior decision-making, path planning, motion control, and other technologies, and has broad application prospects.
[0063] 5) 3D reconstruction: The data process and computer technology that uses 2D projection or images to restore the 3D information of an object.
[0064] 6) Vectorization: Representing an image using geometric primitives based on fundamental data equations such as points, lines, or polygons from computer graphics.
[0065] 7) Affine transformation: In geometry, this refers to the linear transformation of a vector space followed by a translation, which transforms it into another vector space.
[0066] 8) Polygon mesh: In 3D computer graphics, it is a collection of vertices and polygons that represent the shape of a polyhedron, also called an unstructured mesh.
[0067] 9) Intelligent Traffic System (ITS), also known as Intelligent Transportation System, effectively integrates advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. It strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment, and saves energy. Alternatively;
[0068] Intelligent Vehicle Infrastructure Cooperative Systems (IVICS) are a development direction of Intelligent Transportation Systems (ITS). IVICS utilizes advanced wireless communication and next-generation Internet technologies to implement comprehensive, real-time dynamic information exchange between vehicles and infrastructure. Based on the collection and fusion of dynamic traffic information across all times and spaces, it conducts active vehicle safety control and cooperative road management, fully realizing effective collaboration between people, vehicles, and roads. This ensures traffic safety, improves traffic efficiency, and ultimately forms a safe, efficient, and environmentally friendly road traffic system.
[0069] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, and intelligent transportation. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0070] The solutions provided in this application relate to technologies such as computer vision, intelligent transportation, and vehicle-mounted systems in artificial intelligence. These are specifically illustrated through the following embodiments:
[0071] Currently, navigation images typically include, for example: Figure 1 a The vector image shown, and as... Figure 1 b The image shown is a real-world view. For complex intersections with multiple branches, overpasses, or tall buildings obscuring part of the road, vector maps are not clear and intuitive enough for navigation. Therefore, using real-world images to replace vector maps for navigation is gradually becoming the mainstream approach. However, the technology for generating real-world images requires site selection and modeling for each specific intersection. Even if the perspective can be changed, it cannot be reused for other intersections, resulting in high time and manpower costs and reducing the efficiency of real-world image generation.
[0072] This application provides an image processing method, apparatus, device, computer program, and storage medium, which can improve the efficiency of real-scene image generation. The following describes exemplary applications of the electronic devices provided in this application. These electronic devices can be implemented as smartphones, smartwatches, laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), smart voice interaction devices, smart home appliances, and in-vehicle terminals, etc., and can also be implemented as servers. The following describes exemplary applications when the electronic device is implemented as a server.
[0073] See Figure 2 , Figure 2 This is an optional architecture diagram of the image processing system 100 provided in the embodiments of this application. The terminal 400 is connected to the server 200 through the network 300, which can be a wide area network or a local area network, or a combination of the two.
[0074] Server 200 is used to extract features and perform semantic recognition on the road image to be processed to obtain the area to be reconstructed in the road image and the traffic element images in the area to be reconstructed; to perform three-dimensional reconstruction on the area to be reconstructed to obtain a three-dimensional scene point cloud; and to map the traffic element images onto the three-dimensional scene point cloud according to the type of traffic element images to generate a three-dimensional real scene image.
[0075] The server 200 is also used to send the generated 3D real-world image to the terminal 400 for display in the navigation or map APP client 410 installed on the terminal 400.
[0076] In some embodiments, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, or in-vehicle terminal, but is not limited to these.
[0077] In some embodiments, the terminal 400 can also perform feature extraction and semantic recognition on the road image to be processed to obtain the area to be reconstructed in the road image to be processed, and the traffic element image in the area to be reconstructed; perform three-dimensional reconstruction on the area to be reconstructed to obtain a three-dimensional scene point cloud; and map the traffic element image to the three-dimensional scene point cloud according to the type of traffic element image to generate a three-dimensional real scene image and display it on the client 410. The specific selection is made according to the actual situation, and this application embodiment does not limit it.
[0078] The following will describe an exemplary application of an electronic device as a server.
[0079] See Figure 3 , Figure 3 This is a schematic diagram of the structure of the server 200 provided in the embodiments of this application. Figure 2 The server 200 shown includes at least one processor 210, memory 250, at least one network interface 220, and a user interface 230. The various components in server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 240.
[0080] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0081] User interface 230 includes one or more output devices 231 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0082] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.
[0083] The memory 250 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.
[0084] In some embodiments, memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0085] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0086] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0087] Presentation module 253 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with user interface 230 (e.g., a display screen, a speaker, etc.).
[0088] The input processing module 254 is used to detect and translate one or more user inputs or interactions from one or more input devices 232.
[0089] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 3 An image processing device 255 stored in memory 250 is shown. It may be software in the form of programs and plug-ins, including the following software modules: semantic recognition module 2551, three-dimensional reconstruction module 2552 and generation module 2553. These modules are logically related and can therefore be arbitrarily combined or further split according to the functions they implement.
[0090] The functions of each module will be explained below.
[0091] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0092] In some embodiments, the terminal or server can implement the image processing method provided in this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a navigation APP or a map APP; it can also be a mini-program, that is, a program that only needs to be downloaded to a browser environment to run; or it can be a mini-program or web client program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.
[0093] The image processing method provided in this application will be described in conjunction with exemplary applications and implementations of the electronic devices provided in the embodiments of this application.
[0094] See Figure 4 , Figure 4 This is an optional flowchart illustrating an image processing method provided in an embodiment of this application, which will be combined with... Figure 4 The steps shown are explained.
[0095] S101. By performing feature extraction and semantic recognition on the road image to be processed, the area to be reconstructed in the road image to be processed, and the traffic element images in the area to be reconstructed are obtained.
[0096] The image processing method provided in this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, autonomous driving, assisted driving, and vehicle-mounted systems. For example, it can be applied to scenarios where road navigation is performed using navigation software or systems on mobile phones or in-vehicle devices. The specific application can be selected based on actual circumstances, and this application does not impose any limitations.
[0097] In this embodiment, the electronic device can acquire images of the real road environment, such as intersections and lanes, captured by an image acquisition device as the road image to be processed. Since the real road environment captured by the image acquisition device may include images of vehicles or pedestrians on the road, or, based on the acquisition angle of the image acquisition device, images captured by a vehicle dashcam may include images of the vehicle interior, the electronic device can perform feature extraction and semantic recognition on the road image to be processed. This allows it to identify and remove non-road elements such as vehicles, pedestrians, and vehicle interiors from the road image to reduce occlusion of traffic elements, avoid introducing noise into the 3D reconstruction, and obtain the area to be reconstructed. Furthermore, the electronic device can obtain images of traffic elements within the area to be reconstructed based on the semantic recognition of the area to be reconstructed.
[0098] In some embodiments, the electronic device can identify road elements and non-road elements from the road image to be processed through a semantic segmentation network. Other semantic recognition methods can also be used, and the specific method is selected according to the actual situation. This application does not limit the specific method.
[0099] In some embodiments, electronic devices can acquire crowdsourced images of the current road scene as road images to be processed. It is understood that using crowdsourced images to generate real-world images of road scenes can further reduce the cost of manually surveying and mapping road images in related technologies, thereby further improving the efficiency of generating real-world images.
[0100] In this embodiment, the area to be reconstructed is an image region in the road image to be processed from which non-road elements have been removed. The traffic element image is an image in the area to be reconstructed that contains traffic indication information. In some embodiments, the traffic element image may include ground traffic element images, such as ground lane lines, vehicle signals (ground direction indicators), and above-ground traffic element images, such as traffic lights, electronic eyes, speed limit signs, warning signs, road name signs, etc. The specific selection is based on the actual situation, and this embodiment does not limit it.
[0101] S102. Perform 3D reconstruction on the area to be reconstructed to obtain a 3D scene point cloud.
[0102] In this embodiment of the application, the electronic device can perform three-dimensional reconstruction of the area to be reconstructed to obtain a three-dimensional scene point cloud image.
[0103] In some embodiments, the electronic device can recover the three-dimensional geometric structure of the road scene from a two-dimensional image through a camera calibration calculation process, and use computer vision algorithms, such as OpenCV, to calculate the color of each point on the visible surface of the three-dimensional geometric structure based on grayscale, thereby rendering a three-dimensional scene point cloud. Other methods can also be used for three-dimensional reconstruction, and the specific method should be selected according to the actual situation. This application embodiment does not limit the specific method.
[0104] S103. Based on the type of traffic element image, map the traffic element image onto the 3D scene point cloud to generate a 3D real-world image.
[0105] In this embodiment of the application, the electronic device is pre-set with at least one type of traffic element image. For example, the type of traffic element image can be divided based on the universality of traffic instruction information, including general type and specific type.
[0106] For example, general types of traffic element images may include: ground traffic elements (including lane lines, vehicle signals, etc.) and traffic lights, electronic eyes, speed limit signs, warning signs, etc. Specific types of traffic element images may include: road name signs, etc. Electronic devices can map traffic element images onto a 3D scene point cloud using appropriate mapping methods based on the type of traffic element image. In this way, a 3D reality map can be generated without complex modeling or drawing processes.
[0107] In some embodiments, for general-type traffic element images, the electronic device can obtain a template image corresponding to the traffic element image from a preset general vector image library, project the template image onto the three-dimensional scene point cloud, and generate a three-dimensional real-scene image.
[0108] In some embodiments, for specific types of traffic element images, such as planar elements with specific road information like road name signs, it is generally difficult to prepare corresponding vectorized image templates in advance for replacement. Electronic devices can first generate an initial three-dimensional real-scene image based on a three-dimensional point cloud map, and then directly project the specific type of traffic element image onto the initial three-dimensional real-scene image through affine transformation to replace the image of the original point cloud region in the initial three-dimensional real-scene image, thereby obtaining a three-dimensional real-scene image.
[0109] In this embodiment of the application, during the process of reconstructing a three-dimensional scene and generating a three-dimensional scene point cloud in S102, the electronic device can calculate the imaging parameters of the camera, such as the camera pose and the camera intrinsic parameter matrix. The electronic device can use the imaging parameters calculated during the three-dimensional scene reconstruction process to calculate the positional mapping relationship between the traffic element image in the area to be reconstructed and the three-dimensional scene point cloud, thereby realizing the mapping of the traffic element image to the three-dimensional scene point cloud and rendering to generate a three-dimensional real scene image.
[0110] Understandably, by performing semantic recognition on the road image to be processed, the region to be reconstructed in the road image can be obtained, and then a real-scene image can be generated based on the region to be reconstructed, reducing the workload of image processing. Furthermore, according to the type of traffic element image, the traffic element image is mapped to the 3D scene point cloud obtained from the 3D reconstruction of the region to be reconstructed. Compared to related technologies that draw and model each element in the road image, this simplifies the real-scene image generation process, improves processing speed, and thus increases the efficiency of real-scene image generation.
[0111] In some embodiments, based on Figure 4 ,like Figure 5 As shown, before S101, the electronic device can also obtain the road image to be processed by executing S001-S002 through crowdsourced images. The steps will be explained in conjunction with each step.
[0112] S001. Obtain the target image of the current road scene and obtain the current location information corresponding to the current road scene.
[0113] In this embodiment of the application, when the electronic device is a terminal, it can acquire images of the current road scene using an image acquisition device to obtain a target image. The electronic device can also obtain the location information corresponding to the current road scene using a positioning device or positioning module, which serves as its current location information.
[0114] For example, a vehicle traveling on the road can use a camera on its dashcam to capture video or images of its journey as target images. Furthermore, it can obtain the Global Positioning System (GPS) location information of the current image capture location as its current location information.
[0115] In this embodiment of the application, when the electronic device is a server, the electronic device can receive a target image of the current road scene and current location information sent by a first terminal connected to it. Here, the first terminal can be a terminal connected to the server, for which the server provides navigation services for the current road scene.
[0116] In some embodiments, when an electronic device acquires a target image and current location information, it can bind or associate the target image and current location information. For example, the target image can be labeled using the current location information to record its location information. In this way, when summarizing images collected by multiple terminals, multiple images corresponding to the same location can be determined based on the image's location information for subsequent image fusion processing.
[0117] S002. Based on the current location information, acquire at least one crowdsourced image and fuse the at least one crowdsourced image with the target image to obtain the road image to be processed corresponding to the current road scene; the at least one crowdsourced image is an image collected and uploaded by at least one terminal of the current road scene.
[0118] In this embodiment, the electronic device can obtain at least one crowdsourced image corresponding to the current location information. Here, the at least one crowdsourced image is an image collected and uploaded by at least one terminal of the current road scene.
[0119] For example, in a crowdsourcing scenario, multiple vehicles can collect road images or videos while driving using devices such as dashcams. These images or videos, labeled with GPS location information, are then sent to a cloud database. Thus, at the same location, such as intersection A, there may be at least one image uploaded by at least one vehicle in the cloud database. Furthermore, electronic devices, having acquired the target image of intersection A and its current location information, can retrieve at least one road scene image of intersection A uploaded by other vehicles (i.e., other terminals) from the cloud database based on the current location information, as at least one crowdsourced image.
[0120] In this embodiment of the application, the electronic device can use the current location information to perform image fusion with at least one crowdsourced image representing the same location, i.e. the current road scene, and the target image to obtain the crowdsourced image corresponding to the current road scene.
[0121] In some embodiments, electronic devices can process unstructured image or video data into structured data, such as vector data, through calibration and AI algorithms, and then aggregate the vector data to generate crowdsourced images. For example, the crowdsourced image can be a 3D image model of the current road scene, and the image fusion algorithm can be a 3D reconstruction (StructureFromMotion, SFM) algorithm, or a multi-view dense reconstruction (MultiView System, MVS) algorithm, etc., selected according to the actual situation; this application embodiment does not limit the choice.
[0122] In some embodiments, due to differences in the hardware and software settings of image acquisition devices, such as image sensors, the target image and the various crowdsourced images may have inconsistent data sources, accuracy, or format standards. Before performing image fusion, the electronic device may first convert the format of the target image and the at least one crowdsourced image, and then perform image fusion processing after converting them to a unified format.
[0123] Understandably, obtaining road images corresponding to the current road scene through crowdsourcing can reduce the workload and cost of obtaining these images through specialized surveying of the road scene, thereby improving the efficiency of generating 3D reality maps from the road images and reducing the cost of generating 3D reality maps. Furthermore, electronic devices can continuously update the image fusion results using constantly updated crowdsourced images, thus continuously improving the accuracy of image processing and consequently the accuracy of the generated magnified reality map.
[0124] In some embodiments, based on Figure 4 ,like Figure 6 As shown, S101 can be implemented through S1011-S1013, which will be explained in conjunction with each step.
[0125] S1011. Extract features from the road image to be processed to obtain an initial feature point set.
[0126] S1012. Based on the initial feature point set, perform semantic recognition and segmentation to obtain the region in the road image to be processed that is semantically represented as road information, and the traffic element image, which is the image region that is semantically represented as traffic instruction information.
[0127] S1013. The area whose semantic representation is road information is taken as the area to be reconstructed.
[0128] In this embodiment, the electronic device extracts features from the road image to be processed to obtain an initial feature point set, and performs semantic recognition based on the initial feature point set to identify image regions with different semantic meanings in the road image to be processed, such as image regions semantically represented as lanes, surrounding buildings, vehicles, pedestrians, traffic signs, etc. The electronic device uses the image regions semantically represented as road information, such as lanes, surrounding buildings, traffic signs, etc., as the regions to be reconstructed, and uses the images in the image regions semantically represented as traffic instruction information, such as vehicle signals, traffic signs, etc., as traffic element images within the regions to be reconstructed.
[0129] In some embodiments, the electronic device can use a machine learning-trained or multi-object detection network to perform feature extraction, semantic recognition, and semantic segmentation on the road image to be processed, identifying target bounding boxes corresponding to image regions with different semantic meanings from the road image to be processed.
[0130] For example, such as Figure 7 As shown, the road image to be processed is a real scene image of an intersection captured by a dashcam. It can be seen that the initial feature point set obtained by the electronic device through feature extraction of the road image to be processed includes road information, such as lanes, surrounding buildings, ground directional arrows, and road name signs, as well as non-road information, such as feature points of vehicles traveling in the lanes. The electronic device further extracts features from these features. Figure 7 The initial feature point set shown is used for semantic recognition and segmentation. This can identify and delineate regions in the road image to be processed that semantically represent road information, as well as image regions semantically represent traffic indication information, such as... Figure 8 The regions of traffic element images corresponding to each target recognition box in the image.
[0131] In some embodiments, the electronic device can simultaneously identify non-road information regions and traffic element images in the image to be processed using a trained semantic segmentation network, and use the regions outside the non-road information regions as the regions to be reconstructed. In some embodiments, the electronic device can also use different semantic segmentation networks trained with different recognition targets to identify non-road information regions and traffic element images in the image to be processed, respectively. The specific selection is based on the actual situation, and this application embodiment does not limit this.
[0132] In some embodiments, the electronic device can also use a semantic segmentation network trained with road elements as the target to filter out regions in the road image to be processed that are semantically represented as non-road information, i.e., filter out non-road elements in the road image to be processed, and obtain the region to be reconstructed. The specific selection is based on the actual situation, and the embodiments of this application are not limited thereto.
[0133] In some embodiments, based onFigure 6 ,like Figure 8 As shown, S102 can be implemented through S1021-S1022, which will be explained in conjunction with each step.
[0134] S1021. Extract the set of reconstruction feature points corresponding to the area to be reconstructed from the road image to be processed.
[0135] In this embodiment, during the 3D reconstruction process, the electronic device should only restore the relevant model of road information to avoid non-road information such as vehicles or people obscuring traffic elements and affecting the navigation experience. Based on the semantically recognized area to be reconstructed, the electronic device extracts the set of reconstruction feature points corresponding to the area to be reconstructed from the road image to be processed.
[0136] In some embodiments, the electronic device can extract the reconstructed feature point set by adding a mask to the area outside the area to be reconstructed. The electronic device can add a mask to the area outside the area to be reconstructed in the road image to be processed, obtaining the image to be reconstructed; then, it can extract features from the image to be reconstructed. Since the non-road information areas in the image to be reconstructed are covered by the mask, the electronic device can obtain the reconstructed feature point set corresponding to the area to be reconstructed. The electronic device can also achieve targeted feature point extraction of the area to be reconstructed through other methods, depending on the actual situation; this application embodiment does not limit this.
[0137] For example, based on Figure 7 Electronic devices can add masks to non-road information areas, such as... Figure 9 As shown, it can be seen that electronic devices use black pixels instead of black pixels. Figure 7 The pixels of the road information area in Central Africa are obscured. Figure 7 Vehicles driving in the middle lane and parked vehicles on the side of the lane were excluded to avoid adding non-road elements to the 3D reconstruction results.
[0138] S1022. Perform 3D reconstruction using the reconstructed feature point set to obtain a 3D scene point cloud.
[0139] In this embodiment, the electronic device performs 3D reconstruction using a reconstructed feature point set, i.e., a feature point set extracted from the unmasked road information region. Based on the reconstructed feature point set, 3D projection calculations are performed to obtain the imaging parameters of the image acquisition device corresponding to the road image to be processed. Here, the electronic device can save these imaging parameters for subsequent mapping processing. Furthermore, based on the imaging parameters, the electronic device projects each reconstructed feature point in the 2D image into 3D space to obtain a 3D scene point cloud.
[0140] Understandably, by using deep learning networks to filter out non-road elements such as vehicles, people, and vehicle interiors in road images to be processed, it is possible to reduce the introduction of noise in 3D reconstruction, improve the accuracy of 3D reconstruction, and thus improve the accuracy of generating 3D real-world images. It is also possible to reduce the workload of image processing in 3D reconstruction by electronic devices and improve the speed of image processing.
[0141] In some embodiments, based on Figure 4 , Figure 6 and Figure 8 Any one of them, such as Figure 10 As shown, S103 can be implemented through S1031-S1033, which will be explained in conjunction with each step.
[0142] S1031. Obtain the imaging parameters corresponding to the 3D scene point cloud and calculate the imaging trajectory; and based on the imaging trajectory, calculate the positional mapping relationship between the traffic element image and the 3D scene point cloud; the imaging parameters are calculated through the 3D reconstruction process.
[0143] In this embodiment, the electronic device can calculate the imaging parameters corresponding to the 3D scene point cloud through the above-described 3D reconstruction process. In some embodiments, the imaging parameters may include: pose information and intrinsic parameter information of the image acquisition device. Here, the image acquisition device refers to the image acquisition device corresponding to the road image to be processed.
[0144] In this embodiment, the electronic device calculates the camera trajectory of the image acquisition device based on the imaging parameters, which is used as the imaging trajectory. Based on the imaging trajectory, it further calculates the coordinate correspondence between the traffic element image in the two-dimensional space and the three-dimensional scene point cloud, which is used as the positional mapping relationship between the traffic element image and the three-dimensional scene point cloud.
[0145] In some embodiments, the intrinsic parameters include focal length, projection center offset, distortion coefficients, etc. The electronic device can use the imaging parameters containing the intrinsic parameters, along with the Structure from Motion (SFM) algorithm, to calculate the camera's translation (3D vector) and rotation (quaternion) values. Then, using the Bundle Adjustment (BED) algorithm, it can optimize the feature point projection residuals to obtain the imaging trajectory. Based on this trajectory, the positional mapping between the traffic element image and the 3D scene point cloud is calculated.
[0146] S1032. Based on the position mapping relationship, determine the projection area of the traffic element image in the 3D scene point cloud.
[0147] In this embodiment, the electronic device can obtain the three-dimensional coordinates of each two-dimensional pixel in the traffic element image in the three-dimensional scene point cloud according to the position mapping relationship, thereby determining the projection area of the traffic element image in the three-dimensional scene point cloud.
[0148] S1033. Based on the type of traffic element image, map the traffic element image onto the projection area to generate a three-dimensional real-scene image.
[0149] In this embodiment, the electronic device can use a corresponding image mapping method to map the traffic element image onto the projection area according to the type of the traffic element image, thereby generating a three-dimensional real-scene image.
[0150] In some embodiments, based on Figure 10 ,like Figure 11 As shown, for general-type traffic element images, the electronic device can implement the method in S1033 by executing the processes S301-S302, which will be explained in conjunction with each step.
[0151] S301. When the traffic element image is of a general type, obtain the vectorized image corresponding to the traffic element image.
[0152] In some embodiments, based on Figure 11 ,like Figure 12 As shown, S301 can be implemented through S3011 or S3012, as follows:
[0153] S3011. When the traffic element image is of a general type, obtain the target vector template image corresponding to the traffic element image from the preset vector image library and use it as a vectorized image.
[0154] In this embodiment, the electronic device can acquire or access a preset vector image library, wherein the preset vector image library contains at least one vector template image corresponding to at least one general type of traffic element image. The electronic device can obtain the target vector template image corresponding to the traffic element image as a vectorized image by recognizing the image content of the traffic element image and matching the image content with at least one vector template image.
[0155] For example, the traffic element image identified by the electronic device from the road image to be processed can be as follows: Figure 13 As shown, it can be seen that Figure 13 It includes the identified lane guidance images of the ground, such as the images in object detection boxes 12-1 to 12-5. The preset vector library contains vector templates corresponding to various lane guidance images, such as... Figure 14a- Figure 14cAs shown, the electronic device can match the corresponding vector template image in a preset vector library based on each identified lane guidance image, such as the vector template corresponding to the lane guide image in 12-2. Figure 14a , as the target vector template image.
[0156] In some embodiments, the electronic device is based on Figure 13 3D reconstruction yields a 3D scene point cloud, which can be used as follows: Figure 15 As shown, Figure 15 The points in the image represent feature point clouds, the cone-shaped bounding box represents the imaging trajectory, and the remaining two-dimensional rectangular boxes represent the position boxes of the identified traffic element images in three-dimensional space.
[0157] S3012. When the traffic element image is of a general type, replace the traffic element image with a grid, and perform vectorization drawing based on the grid to obtain a vectorized image.
[0158] In this embodiment, the 3D scene point cloud obtained from 3D reconstruction is usually an unstructured sparse point cloud. Directly generating a real-world image based on the 3D scene point cloud may not intuitively provide users with a clear and accurate location. Dense reconstruction provides better scene restoration and representation, but it is time-consuming and requires a large amount of computing resources. To improve processing efficiency, when the traffic element image is of a general type, the electronic device can also combine prior knowledge, such as the fact that all ground or lane lines are rectangular, to replace the traffic element image with a polygon mesh. The meshed traffic element image is then placed in a virtual 3D space and drawn from a specific viewpoint to obtain a vectorized image.
[0159] S302. Replace the projection area with a vectorized image to update the 3D scene point cloud, and render the updated 3D scene point cloud based on a preset angle to obtain a 3D real-world image.
[0160] In this embodiment, the electronic device replaces the projected area in the 3D scene point cloud with a vectorized image, thereby updating the 3D scene point cloud and obtaining an updated 3D scene point cloud. Then, based on a preset angle, the updated 3D scene point cloud is rotated and rendered to obtain a 3D real-world image.
[0161] In some embodiments, the electronic device may also implement the method in S1033 by executing S303-S304, which will be described in conjunction with each step.
[0162] S303. When the traffic element image is of a general type, extract the image texture from the traffic element image, add the image texture to the projection area, update the 3D scene point cloud, and obtain the updated 3D scene point cloud.
[0163] S304. Based on a preset angle, render the updated 3D scene point cloud to obtain a 3D real-world image.
[0164] In some embodiments, the three-dimensional real-world image generated by the electronic device through the vectorized image replacement process in S301-S302 or S303-S304 can be as follows: Figure 16 As shown.
[0165] It is understood that in the embodiments of this application, when the traffic element image is of a general type, replacing the two-dimensional traffic element image frame obtained through semantic recognition with basic vectorized elements, or directly extracting image textures for addition, can save the workload of modeling and drawing each element in the original image during the real-scene image generation process and improve the generation efficiency of the real-scene image.
[0166] In some embodiments, based on Figure 10 ,like Figure 17 As shown, for a specific type of traffic element image, the electronic device can implement the method in S1033 by executing the processes S305-S306, which will be explained in conjunction with each step.
[0167] S305. When the traffic element image is of a specific type, render the three-dimensional scene point cloud or the updated three-dimensional scene point cloud based on a preset angle to generate an initial three-dimensional real scene image.
[0168] In this embodiment, when the traffic element image is of a specific type, it is often difficult to replace it with a pre-prepared vectorized template image. For example, the text content in a road name sign image, such as "Road X," cannot usually be replaced by a template. Furthermore, the clarity of the 3D real-world image generated solely by rendering the point cloud portion corresponding to the road name sign image in the 3D scene point cloud may not meet navigation requirements. Figure 16 In this case, the image of the road name sign 15-1, rendered solely from sparse point clouds, is rather blurry and cannot effectively indicate the road name. Therefore, the electronic device can first render the 3D scene point cloud or an updated 3D scene point cloud based on a preset angle to generate an initial 3D real-world image, and then update the image portions corresponding to specific types of traffic elements in the initial 3D real-world image.
[0169] Here, when the area to be reconstructed contains both general-type traffic element images and specific-type traffic element images, the electronic device can first update the region corresponding to the general-type traffic element image in the 3D scene point cloud using methods in S301-S302 or S303, obtaining an updated 3D scene point cloud. Then, based on the updated 3D scene point cloud and a preset angle, rendering is performed to generate an initial 3D reality image. In this case, the initial 3D reality image already contains the image mapping portion of the general-type traffic element image.
[0170] In some embodiments, when the area to be reconstructed contains only images of specific types of traffic elements, the electronic device can render based on the three-dimensional scene point cloud and a preset angle to generate an initial three-dimensional real-world image.
[0171] S306. Based on the location information of the projection area in the initial three-dimensional real scene map, update the initial three-dimensional real scene map using traffic element images to obtain the three-dimensional real scene map.
[0172] In this embodiment, the electronic device can obtain the position information of the projection area in the initial three-dimensional real scene map based on the projection area in the three-dimensional scene point cloud. The electronic device directly uses the original two-dimensional image of the traffic element image to update the initial three-dimensional real scene map, thereby realizing the affine transformation of the traffic element image to obtain the three-dimensional real scene map.
[0173] For example, images of specific types of traffic elements can be as follows: Figure 13 Road name signs 12-6, or Figure 18 The road name sign 17-1 is shown in the figure. The electronic device can obtain the four corner points of the target detection box corresponding to the road name sign, and obtain the position information of the projection area in the initial three-dimensional real scene image through formula (1). Then, through the method in S304-S305, the road name sign 17-1 is subjected to affine transformation to obtain the corresponding image part 17-2 in the three-dimensional real scene image. Formula (1) is as follows:
[0174] (1)
[0175] In formula (1), This refers to images of traffic elements, such as the two-dimensional coordinate information of road name signs. The positional mapping relationship between the two-dimensional image and the initial three-dimensional real-world image, that is, the transformation relationship between the two-dimensional image coordinate system and the three-dimensional spatial coordinate system. for The corresponding 3D coordinate information in the initial 3D reality image.
[0176] It is understood that, in the embodiments of this application, when the traffic element image is of a specific type, the electronic device can use affine transformation to replace the blurred image generated by reconstruction with the original traffic element image, such as a road name sign. This not only improves the image accuracy of the 3D real-scene image (as can be seen), Figure 18 The clarity of the 17-2 is significantly higher than that of the others. Figure 16 The process of generating 3D reality images (15-1) has been simplified, the rendering workload has been reduced, and the efficiency of generating 3D reality images has been improved.
[0177] In some embodiments, the traffic element image may further include: at least one road line detection point corresponding to at least one road line, wherein each road detection point among the at least one road line detection point may contain at least one road detection point. For example... Figure 13 As shown, the electronic device can obtain at least one road line detection point 12-7 corresponding to the double yellow line detection point among at least one type of road line detection points through a semantic recognition process. Similarly, it can also obtain at least one road line detection point corresponding to other types of lane lines. For each type of road line detection point among at least one type of road line detection points, the sub-device can perform three-dimensional projection and topological connection based on at least one road line detection point of that type to vectorize the detection point results of the two-dimensional image and generate a road line image in the three-dimensional real scene.
[0178] The following will combine Figure 19 This illustrates an exemplary application of the embodiments of this application in a real-world scenario.
[0179] This application embodiment can be applied in mobile phone navigation software. Through S401-S40, a real-view magnified image of the intersection is generated to provide road navigation to the user, as follows.
[0180] S401, Obtain crowdsourced images.
[0181] In S401, the electronic device acquires crowdsourced images as road images to be processed.
[0182] S402, Semantic Segmentation.
[0183] In S402, the electronic device performs semantic segmentation on the crowdsourced image, identifying elements such as vehicles and people (equivalent to non-road information areas), basic traffic elements (equivalent to general-type traffic element images), and road name signs (equivalent to specific-type traffic element images). The execution process of S402 is consistent with the description in S1011-S1012 above, and will not be repeated here.
[0184] S403, Add masking and semantics.
[0185] In S403, the electronic device adds masks to the identified vehicle, pedestrian and other elements to obtain a non-road information area with added masks, and the basic traffic elements and road name signs have semantic information to be reconstructed.
[0186] S404, 3D reconstruction.
[0187] In S404, the electronic device performs feature extraction and 3D reconstruction on the image to be reconstructed, obtaining a 3D scene point cloud. The execution process of S403-S404 is consistent with the description of S1021-S1022 above, and will not be repeated here.
[0188] S405. The basic traffic elements are replaced by vectors to update the 3D scene point cloud, resulting in an updated 3D scene point cloud. The updated 3D scene point cloud is then rendered to obtain an initial 3D real-world image.
[0189] The execution process of S405 is the same as described above. Figure 10 The descriptions of S1031-S301 are consistent and will not be repeated here.
[0190] S406. Perform an affine transformation on the road name sign and replace the corresponding area of the initial 3D real-scene image to obtain a 3D real-scene image.
[0191] The execution process of S405 is the same as described above. Figure 15 The descriptions of S1031-S3061 are consistent and will not be repeated here.
[0192] It is understood that in this embodiment, only crowdsourced data is used to reconstruct the road environment, resulting in a real-scene intersection magnification image that can be used for navigation. This significantly reduces the cost of real-scene image production and eliminates the need for manual image collection, drawing, and modeling of the road. It greatly simplifies the existing production process for real-scene intersection magnification images. The solution in this embodiment is low-cost, easily replicable, and can significantly improve the speed of generating real-scene intersection magnification images.
[0193] The following description continues to illustrate the exemplary structure of the image processing apparatus 255 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 3 As shown, the software modules stored in the image processing device 255 of the memory 250 may include:
[0194] The semantic recognition module 2551 is used to extract features and perform semantic recognition on the road image to be processed to obtain the area to be reconstructed in the road image to be processed, and the traffic element image in the area to be reconstructed.
[0195] The 3D reconstruction module 2552 is used to perform 3D reconstruction on the area to be reconstructed to obtain a 3D scene point cloud.
[0196] The generation module 2553 is used to map the traffic element image onto the three-dimensional scene point cloud according to the type of the traffic element image to generate a three-dimensional real scene image.
[0197] In the above-mentioned device, the semantic recognition module 2551 is further configured to extract features from the road image to be processed to obtain an initial feature point set; perform semantic recognition and segmentation based on the initial feature point set to obtain the region in the road image to be processed that is semantically represented as road information, and the traffic element image, wherein the traffic element image is an image region that is semantically represented as traffic indication information; and use the region that is semantically represented as road information as the region to be reconstructed.
[0198] In the above-mentioned device, the three-dimensional reconstruction module 2552 is further configured to extract the set of reconstruction feature points corresponding to the area to be reconstructed from the road image to be processed; and to perform three-dimensional reconstruction using the set of reconstruction feature points to obtain the three-dimensional scene point cloud.
[0199] In the above-mentioned device, the three-dimensional reconstruction module 2552 is further configured to add a mask to the area outside the area to be reconstructed in the road image to be processed, so as to obtain the image to be reconstructed; and to extract features from the image to be reconstructed to obtain the reconstructed feature point set.
[0200] In the above-described device, the image processing device 255 further includes an acquisition module and a fusion module. The acquisition module is used to acquire a target image of the current road scene and acquire the current location information corresponding to the current road scene before performing feature extraction and semantic recognition on the road image to be processed; based on the current location information, acquire at least one crowdsourced image and fuse the at least one crowdsourced image with the target image to obtain the road image to be processed corresponding to the current road scene; the at least one crowdsourced image is an image collected and uploaded by at least one terminal of the current road scene.
[0201] In the above-described device, the generation module 2553 is used to acquire the imaging parameters corresponding to the three-dimensional scene point cloud, calculate the imaging trajectory, and calculate the positional mapping relationship between the traffic element image and the three-dimensional scene point cloud based on the imaging trajectory; the imaging parameters are calculated through a three-dimensional reconstruction process; the projection area corresponding to the traffic element image in the three-dimensional scene point cloud is determined according to the positional mapping relationship; and the traffic element image is mapped to the projection area according to the type of the traffic element image to generate the three-dimensional real scene image.
[0202] In the above-mentioned device, the type of the traffic element image includes: general type. The generation module 2553 is further used to obtain the vectorized image corresponding to the traffic element image when the traffic element image is of general type; replace the projection area with the vectorized image to update the three-dimensional scene point cloud, and render the updated three-dimensional scene point cloud based on a preset angle to obtain the three-dimensional real scene image.
[0203] In the above-described apparatus, the generation module 2553 is further configured to obtain a target vector template image corresponding to the traffic element image from a preset vector image library, and use it as the vectorized image; the preset vector image library contains at least one vector template image corresponding to at least one general type of traffic element image; or, replace the traffic element image with a grid, and perform vectorized drawing based on the grid to obtain the vectorized image.
[0204] In the above-mentioned device, the generation module 2553 is further configured to extract image texture from the traffic element image when the traffic element image is of a general type, add the image texture to the projection area, update the three-dimensional scene point cloud, and obtain an updated three-dimensional scene point cloud; and render the updated three-dimensional scene point cloud based on a preset angle to obtain the three-dimensional real scene image.
[0205] In the above-mentioned device, the type of traffic element image includes: a specific type. The generation module 2553 is further configured to, when the traffic element image is of a specific type, render the three-dimensional scene point cloud or the updated three-dimensional scene point cloud based on a preset angle to generate an initial three-dimensional real scene image; and update the initial three-dimensional real scene image using the traffic element image according to the position information corresponding to the projection area in the initial three-dimensional real scene image to obtain the three-dimensional real scene image.
[0206] In the above-mentioned device, the traffic element image includes at least one road line detection point. The generation module 2553 is further used to perform three-dimensional projection and topological connection on the at least one road line detection point to generate the road line image in the three-dimensional real scene.
[0207] It should be noted that the description of the above device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0208] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to perform the method provided in this application, for example... Figure 4- Figure 6 ,Figure 8 , Figure 10- Figure 12 , Figure 17 or Figure 19 The method shown in the figure.
[0209] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0210] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0211] As an example, executable instructions may, but do not necessarily, correspond to files in the file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0212] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0213] In summary, semantic recognition of the road image to be processed can identify the region to be reconstructed, allowing for the generation of a real-scene image based on this region, thus reducing the workload of image processing. Furthermore, by mapping traffic element images to the 3D point cloud obtained from the 3D reconstruction of the region to be reconstructed, the real-scene image generation process is simplified compared to drawing and modeling each element in the road image in related technologies. This improves processing speed and efficiency. Moreover, by using deep learning networks to filter non-road elements such as vehicles, people, and vehicle interiors in the road image, noise introduction during 3D reconstruction is reduced, improving the accuracy of 3D reconstruction and thus the accuracy of the generated 3D real-scene image. This also reduces the workload of image processing in electronic devices during 3D reconstruction, increasing processing speed. Furthermore, when the traffic element images are of a general type, replacing the 2D traffic element image boxes obtained through semantic recognition with basic vectorized elements, or directly extracting and adding image textures, saves the workload of modeling and drawing each element in the original image during real-scene image generation, further improving generation efficiency. When the traffic element image is of a specific type, electronic devices can use affine transformation to replace the blurry image generated during reconstruction with the original traffic element image, such as a road name sign. This not only improves the image accuracy of the 3D reality image but also increases the generation efficiency of the 3D reality image.
[0214] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An image processing method, characterized by, The method comprises the following steps: obtaining a to-be-reconstructed region in the to-be-processed road image and a traffic element image in the to-be-reconstructed region by performing feature extraction and semantic recognition on the to-be-processed road image; performing three-dimensional reconstruction on the to-be-reconstructed region to obtain a three-dimensional scene point cloud; obtaining imaging parameters corresponding to the three-dimensional scene point cloud, calculating an imaging track, and calculating a position mapping relationship between the traffic element image and the three-dimensional scene point cloud based on the imaging track; the imaging parameters are calculated through the three-dimensional reconstruction process; determining a projection region corresponding to the traffic element image in the three-dimensional scene point cloud according to the position mapping relationship; mapping the traffic element image into the projection region according to the type of the traffic element image to generate a three-dimensional real scene map.
2. The method of claim 1, wherein, The method comprises the following steps: performing feature extraction on the to-be-processed road image to obtain an initial feature point set; performing semantic recognition and segmentation based on the initial feature point set to obtain a region in the to-be-processed road image that is semantically represented as road information and the traffic element image that is an image region semantically represented as traffic indication information; regarding the region semantically represented as road information as the to-be-reconstructed region.
3. The method of claim 2, wherein, The method comprises the following steps: extracting a reconstruction feature point set corresponding to the to-be-reconstructed region from the to-be-processed road image; performing three-dimensional reconstruction using the reconstruction feature point set to obtain the three-dimensional scene point cloud.
4. The method of claim 3, wherein, The method comprises the following steps: adding a mask to a region outside the to-be-reconstructed region in the to-be-processed road image to obtain a to-be-reconstructed image; performing feature extraction on the to-be-reconstructed image to obtain the reconstruction feature point set.
5. The method of claim 1, wherein, Before the feature extraction and semantic recognition of the to-be-processed road image, the method further comprises the following steps: obtaining a target image of a current road scene and obtaining current positioning information corresponding to the current road scene; obtaining at least one crowd-sourced image according to the current positioning information, and performing image fusion on the at least one crowd-sourced image and the target image to obtain a to-be-processed road image corresponding to the current road scene; the at least one crowd-sourced image is an image collected and uploaded by at least one terminal for the current road scene.
6. The method of claim 1, wherein, The type of the traffic element image includes a general type, and the method of mapping the traffic element image into the projection region according to the type of the traffic element image to generate a three-dimensional real scene map comprises the following steps: in the case that the traffic element image is of the general type, obtaining a vectorized image corresponding to the traffic element image; updating the three-dimensional scene point cloud by replacing the projection region with the vectorized image, and rendering the updated three-dimensional scene point cloud based on a preset angle to obtain a three-dimensional real scene map.
7. The method of claim 6, wherein, The method of obtaining a vectorized image corresponding to the traffic element image comprises the following steps: obtain a target vector template image corresponding to the traffic element image from a preset vector image library as the vectorized image; the preset vector image library contains at least one vector template image corresponding to at least one traffic element image of a general type; Or, replace the traffic element image with a grid, and perform vectorization drawing based on the grid to obtain the vectorized image.
8. The method of claim 1, wherein, The method further comprises: in a case where the traffic element image is of a general type, extracting an image texture from the traffic element image, adding the image texture to the projection area, updating the three-dimensional scene point cloud to obtain an updated three-dimensional scene point cloud; rendering the updated three-dimensional scene point cloud based on a preset angle to obtain a three-dimensional real scene image.
9. The method according to any one of claims 6-8, characterized in that, The type of the traffic element image includes a specific type, and the mapping of the traffic element image into the projection area to generate the three-dimensional real scene image according to the type of the traffic element image comprises: in a case where the traffic element image is of a specific type, rendering the three-dimensional scene point cloud or the updated three-dimensional scene point cloud based on a preset angle to generate an initial three-dimensional real scene image; updating the initial three-dimensional real scene image using the traffic element image according to the position information of the projection area in the initial three-dimensional real scene image to obtain the three-dimensional real scene image.
10. The method according to any one of claims 1 to 5, characterized in that, The traffic element image includes at least one road line detection point, and the method further comprises: performing three-dimensional projection and topological connection on the at least one road line detection point to generate a road line image in the three-dimensional real scene image.
11. An image processing apparatus characterized by comprising: comprises: a semantic recognition module configured to obtain a region to be reconstructed in a to-be-processed road image and a traffic element image in the region to be reconstructed by performing feature extraction and semantic recognition on the to-be-processed road image; a three-dimensional reconstruction module configured to perform three-dimensional reconstruction on the region to be reconstructed to obtain a three-dimensional scene point cloud; a generation module configured to obtain imaging parameters corresponding to the three-dimensional scene point cloud, calculate an imaging track, and calculate a position mapping relationship between the traffic element image and the three-dimensional scene point cloud based on the imaging track; the imaging parameters are calculated through a three-dimensional reconstruction process; determine a projection area corresponding to the traffic element image in the three-dimensional scene point cloud according to the position mapping relationship; map the traffic element image into the projection area according to the type of the traffic element image to generate a three-dimensional real scene image.
12. The apparatus of claim 11, wherein, The semantic recognition module is further configured to: perform feature extraction on the to-be-processed road image to obtain an initial feature point set, perform semantic recognition and segmentation based on the initial feature point set to obtain a region in the to-be-processed road image whose semantic representation is road information and the traffic element image, and the traffic element image is an image region whose semantic representation is traffic indication information; and take the region whose semantic representation is road information as the region to be reconstructed.
13. The apparatus of claim 12, wherein, The three-dimensional reconstruction module is further configured to: extract a reconstruction feature point set corresponding to the region to be reconstructed from the to-be-processed road image; and perform three-dimensional reconstruction using the reconstruction feature point set to obtain the three-dimensional scene point cloud.
14. The apparatus of claim 13, wherein, The three-dimensional reconstruction module is further configured to: In the to-be-processed road image, a mask is added to a region outside the to-be-reconstructed region to obtain a to-be-reconstructed image; and feature extraction is performed on the to-be-reconstructed image to obtain the reconstructed feature point set.
15. The apparatus of claim 11, wherein, Before the feature extraction and semantic recognition on the to-be-processed road image, the method further includes an acquisition module and a fusion module, configured to: acquire a target image of a current road scene and acquire current positioning information corresponding to the current road scene; acquire at least one crowd-sourced image according to the current positioning information, and perform image fusion on the at least one crowd-sourced image and the target image to obtain a to-be-processed road image corresponding to the current road scene; the at least one crowd-sourced image is an image collected and uploaded by at least one terminal on the current road scene.
16. The apparatus of claim 11, wherein, The type of the traffic element image includes a general type, and the generation module is further configured to: in a case where the traffic element image is of the general type, acquire a vectorized image corresponding to the traffic element image; replace the projection region with the vectorized image, update the three-dimensional scene point cloud, and render the updated three-dimensional scene point cloud based on a preset angle to obtain a three-dimensional real scene map.
17. The apparatus of claim 16, wherein, The generation module is further configured to: acquire a target vector template image corresponding to the traffic element image from a preset vector image library as the vectorized image; the preset vector image library includes at least one vector template image corresponding to at least one traffic element image of a general type; or replace the traffic element image with a grid, perform vectorized drawing based on the grid to obtain the vectorized image.
18. An electronic device, comprising: comprise: a memory configured to store executable instructions; a processor configured to execute the executable instructions stored in the memory to implement the method in any one of claims 1 to 10.
19. A computer-readable storage medium, characterized in that, executable instructions stored in the memory, configured to be executed by the processor to implement the method in any one of claims 1 to 10.
20. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the method in any one of claims 1 to 10.
Citation Information
Patent Citations
Image fusion method, automatic driving control method, device and equipment
CN110969592A