Traffic Sign Detection Method and Related Equipment
By directly collecting and processing real-life traffic images on the vehicle terminal and detecting the main and auxiliary traffic signs, the navigation delay and instability caused by abnormal network communication in the intelligent traffic system is solved, and efficient and safe navigation data acquisition is achieved.
Patent Information
- Application Number
- CN202110528690.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-05-14
AI Technical Summary
In the existing intelligent transportation system, network communication between the map service equipment and the vehicle terminal is abnormal, resulting in delayed navigation data or failed reception, reducing the real-time and reliability of navigation, especially posing a threat to the driving safety of autonomous vehicles.
A traffic sign detection method is proposed. Through the vehicle terminal, the real traffic scene images are directly collected, image compression is performed, and the resolution is reduced. The main traffic sign is detected first, and then the main traffic sign is used to locate the potential image area of the auxiliary traffic sign to conduct auxiliary traffic sign detection, avoiding excessive occupation of the vehicle terminal resources.
It realizes rapid and accurate detection of traffic signs without relying on servers, improves the real-time and reliability of navigation, ensures the driving safety of autonomous vehicles, and reduces the use of computing resources.
Smart Images

Figure CN113221756B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent transportation applications, and particularly to a traffic sign detection method and related devices. Background Art
[0002] Nowadays, with the application of intelligent transportation systems in scenarios such as intelligent driving and driverless driving, map service devices usually provide road guidance data for in-vehicle terminals based on the traffic conditions between the starting position and the target position of a vehicle's travel, so as to guide the vehicle to safely and smoothly drive to the destination.
[0003] Taking the map navigation scenario as an example, usually, the in-vehicle terminal connects to the network and sends a navigation request to the map service device, so that the map server responds to the navigation request, determines the corresponding road image collected by the map collector using the image collection device, and feeds it back to the in-vehicle terminal for display to guide the vehicle's travel.
[0004] However, in practical applications, due to various factors such as the external environment or the in-vehicle terminal itself, network communication between the map service device and the in-vehicle terminal is often abnormal, resulting in problems such as data delay or reception failure, thereby reducing the real-time performance and reliability of vehicle travel navigation, and even threatening the driving safety of autonomous vehicles. Summary of the Invention
[0005] In view of this, the present application proposes the following technical solutions:
[0006] On the one hand, the present application proposes a traffic sign detection method, the method comprising:
[0007] Obtain a traffic real-scene image in the vehicle travel direction;
[0008] Perform compression processing on the traffic real-scene image to obtain an image to be detected; the resolution of the image to be detected is less than the resolution of the traffic real-scene image;
[0009] Perform main traffic sign detection on the image to be detected, and based on the main traffic sign detection result, obtain a potential image area where a first auxiliary traffic sign matching the detected first main traffic sign exists; the potential image area includes a local image area of the traffic real-scene image;
[0010] Perform auxiliary traffic sign detection on the potential image area, and based on the auxiliary traffic sign detection result, obtain the target position information of the detected first auxiliary traffic sign mapped on the traffic real-scene image.
[0011] In some embodiments, for the main traffic sign detection of the image to be detected, according to the main traffic sign detection result, obtaining a potential image region where there is a first auxiliary traffic sign that matches the detected first main traffic sign includes:
[0012] Inputting the image to be detected into a main traffic sign detection network to obtain the first main traffic sign included in the traffic scene image;
[0013] Taking the first main traffic sign included in the traffic scene image as a reference object to determine a local image region where there is a first auxiliary traffic sign that matches the first main traffic sign;
[0014] The potential image region is composed of the determined local image regions.
[0015] In some embodiments, taking the first main traffic sign included in the traffic scene image as a reference object to determine a local image region where there is a first auxiliary traffic sign that matches the first main traffic sign includes:
[0016] In the traffic scene image, taking the obtained first main traffic sign as the center and spreading a preset distance around to obtain a local image region where there is a first auxiliary traffic sign that matches the first main traffic sign.
[0017] In some embodiments, during the process of spreading a preset distance around with the obtained first main traffic sign as the center, the method further includes:
[0018] Detecting that the first boundary distance between the first main traffic sign and the first image boundary is less than the preset distance; the first image boundary refers to the boundary of the traffic scene image in the first spreading direction;
[0019] In the first spreading direction, spreading from the first main traffic sign to the first image boundary and continuing to perform image filling and spreading according to the first image filling method until the spreading distance in the first spreading direction reaches the preset distance;
[0020] The potential image region is composed of the determined local image regions, including:
[0021] The potential image region where there is the first auxiliary traffic sign is composed of the local image region and the image filling region in the obtained traffic scene image.
[0022] In some embodiments, taking the first main traffic sign included in the traffic scene image as a reference object to determine a local image region where there is a first auxiliary traffic sign that matches the first main traffic sign includes:
[0023] Obtain the first position deployment relationship corresponding to the main traffic sign of the sign type according to the sign type of the first main traffic sign included in the traffic real-scene image; wherein, the first position deployment relationship refers to the relative position relationship between the main traffic sign of the sign type and the auxiliary traffic sign matching the main traffic sign in traffic road facilities.
[0024] According to the first position deployment relationship, determine, from the traffic real-scene image, a partial image area where there is a first auxiliary traffic sign matching the first main traffic sign.
[0025] In some embodiments, the composed potential image area by the determined partial image areas includes:
[0026] Compare the area of the determined partial image area with a preset auxiliary traffic sign detection area;
[0027] If the area of the partial image area is smaller than the preset auxiliary traffic sign detection area, perform a magnification process on the partial image area to obtain a potential image area with the preset auxiliary traffic sign detection area;
[0028] If the area of the partial image area is larger than the preset auxiliary traffic sign detection area, perform a compression process on the partial image area to obtain a potential image area with the preset auxiliary traffic sign detection area;
[0029] Wherein, the resolution of the potential image area is smaller than the resolution of the traffic real-scene image.
[0030] In some embodiments, for the auxiliary traffic sign detection on the potential image area, according to the auxiliary traffic sign detection result, obtain the target position information of the detected first auxiliary traffic sign mapped on the traffic real-scene image, including:
[0031] Input the potential image area into an auxiliary traffic sign detection network to obtain the first auxiliary traffic sign included in the potential image area and the regional position information of the first auxiliary traffic sign in the potential image area;
[0032] Obtain the potential regional position information of the potential image area in the traffic real-scene image;
[0033] According to the obtaining method of the potential image area and the potential regional position information, perform coordinate conversion processing on the regional position information to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0034] In some embodiments, if the number of the first main traffic signs is multiple, obtaining a potential image area where a first auxiliary traffic sign matching the detected first main traffic sign exists includes:
[0035] According to the detected multiple first main traffic signs, obtaining potential image areas where the first auxiliary traffic signs respectively matching the multiple first main traffic signs exist; wherein, the potential image area includes the corresponding first auxiliary traffic sign and at least part of the first main traffic sign matching the first auxiliary traffic sign;
[0036] Performing auxiliary traffic sign detection on the potential image area includes:
[0037] Performing auxiliary traffic sign detection on each obtained potential image area.
[0038] In some embodiments, performing coordinate conversion processing on the area position information according to the obtaining manner of the potential image area and the potential area position information to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image includes:
[0039] Determining the image filling data of the traffic real-scene image and the scaling ratio of the local image area where the first auxiliary traffic sign exists obtained from the traffic real-scene image during the process of obtaining the potential image area; wherein, the image filling data includes the filling distances in different diffusion directions; the potential image area is obtained by compressing or enlarging the local image area;
[0040] Using the filling distances in different diffusion directions, the scaling ratio and the potential area position data to perform restoration processing on the obtained area position information to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0041] In some embodiments, the method further includes:
[0042] Rendering the traffic real-scene image according to the target position information;
[0043] Outputting the rendered navigation image, and displaying the detected first main traffic sign and the first auxiliary traffic sign matching the first main traffic sign in the navigation image; or,
[0044] Popping up a traffic prompt window in the navigation image, and displaying traffic prompt information for the first auxiliary traffic sign in the traffic prompt window; the traffic prompt information is determined based on the auxiliary traffic indication information of the first auxiliary traffic sign, and the driving state information and / or system time of the vehicle; or,
[0045] Perform voice broadcast of the traffic prompt information.
[0046] In some embodiments, the method further includes:
[0047] Obtain the auxiliary traffic indication information of the corresponding first auxiliary traffic sign according to the target location information;
[0048] Utilize the target location information and the auxiliary traffic indication information to obtain the structured perception result of the corresponding first auxiliary traffic sign in the traffic real-scene image;
[0049] Report the structured perception result to the map service device, and the map service device updates the corresponding map data by using the structured perception result.
[0050] In some embodiments, the method further includes:
[0051] Obtain the relative distance between the vehicle and the corresponding first affiliated traffic sign according to the target location information;
[0052] Obtain the main traffic indication information of the first main traffic sign included in the traffic real-scene image, and the auxiliary traffic indication information of the first auxiliary traffic sign that matches the first main traffic sign;
[0053] Generate a road guidance message in the driving direction of the vehicle by using the relative distance, the main traffic indication information, and the auxiliary traffic indication information;
[0054] Perform voice broadcast of the road guidance message.
[0055] In another aspect, the present application also proposes a traffic sign detection device, and the device includes:
[0056] A traffic real-scene image acquisition module, configured to acquire a traffic real-scene image in the driving direction of the vehicle;
[0057] A to-be-detected image obtaining module, configured to perform compression processing on the traffic real-scene image to obtain a to-be-detected image; the resolution of the to-be-detected image is less than the resolution of the traffic real-scene image;
[0058] A potential image area obtaining module, configured to perform main traffic sign detection on the to-be-processed traffic image, and obtain a potential image area where there is a first auxiliary traffic sign that matches the detected first main traffic sign according to the main traffic sign detection result; the potential image area includes a local image area of the traffic real-scene image;
[0059] A target position information acquisition module, configured to perform auxiliary traffic sign detection on the potential image region, and obtain the target position information of the detected first auxiliary traffic sign mapped on the traffic real-scene image according to the auxiliary traffic sign detection result.
[0060] In another aspect, the present application also provides a terminal device, which includes:
[0061] A communication interface;
[0062] A memory, configured to store a program for implementing the traffic sign detection method as described above;
[0063] A processor, configured to load and execute the program stored in the memory to implement the steps of the traffic sign detection method as described above.
[0064] In another aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is called and executed by a processor to implement the traffic sign detection method as described above.
[0065] Based on the above technical solutions, the present application proposes to directly collect the traffic real-scene image in the driving direction of the vehicle by the in-vehicle terminal, and identify the auxiliary traffic signs included therein by the local in-vehicle terminal. The entire detection process does not rely on a server, avoiding various problems caused by poor communication network. Moreover, in order to be applicable to in-vehicle terminals with relatively poor computing capabilities, the present application compresses the traffic real-scene image collected in real time, reduces the image resolution, and performs main traffic sign detection on the obtained low-resolution image to be detected, and locates the first main traffic sign in the traffic real-scene image. Compared with the processing method of directly performing traffic sign detection on the traffic real-scene image, the resource occupancy of the in-vehicle terminal is greatly reduced, while reliably detecting the main traffic sign with a large size, avoiding lags or crashes during the processing process and affecting the detection efficiency.
[0066] After that, the in-vehicle terminal will obtain the potential image region of the first auxiliary traffic sign that exists in the traffic real-scene image and matches the detected first main traffic sign based on the prior knowledge of the main traffic sign detection result, and then perform auxiliary traffic sign detection on it, and can quickly and accurately obtain the auxiliary traffic sign detection result, that is, obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image. It can be seen that compared with the processing method of directly detecting auxiliary traffic signs from the traffic real-scene image collected locally, the present application greatly reduces the calculation amount in the entire traffic sign detection process while ensuring the accurate positioning of the auxiliary traffic signs in the traffic real-scene image, ensuring that in-vehicle terminals with relatively poor local computing capabilities are sufficient to support implementation, and meeting the accurate detection requirements of auxiliary traffic signs in different application scenarios. Description of the Drawings
[0067] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0068] Figure 1 It shows a schematic diagram of the hardware structure of an optional example of a terminal device for implementing the traffic sign detection method provided by the present application;
[0069] Figure 2 It shows a schematic diagram of the hardware structure of another optional example of a terminal device for implementing the traffic sign detection method provided by the present application;
[0070] Figure 3 It shows a schematic diagram of the process of an optional example of the traffic sign detection method proposed by the present application;
[0071] Figure 4 It shows a schematic diagram of the positional relationship between the main traffic sign and its matching auxiliary traffic sign in the traffic sign detection method proposed by the present application;
[0072] Figure 5 It shows a schematic diagram of the process of another optional example of the traffic sign detection method proposed by the present application;
[0073] Figure 6 It shows a schematic diagram of the process of an optional implementation for obtaining a potential image area from a traffic real-scene image in the traffic sign detection method proposed by the present application;
[0074] Figure 7 It shows a schematic diagram of the process of another optional implementation for obtaining a potential image area from a traffic real-scene image in the traffic sign detection method proposed by the present application;
[0075] Figure 8 It shows a schematic diagram of the application process of another optional example of the traffic sign detection method proposed by the present application;
[0076] Figure 9 It shows a schematic diagram of the structure of an optional example of the traffic sign detection device proposed by the present application;
[0077] Figure 10 It shows a schematic diagram of the structure of another optional example of the traffic sign detection device proposed by the present application;
[0078] Figure 11 It shows a schematic diagram of the system architecture of an optional application environment applicable to the traffic sign detection method and device proposed by the present application. Detailed implementation mode
[0079] In view of the technical problems described in the background art section, the present application proposes that a terminal device on the vehicle (i.e., an in-vehicle terminal) be used to achieve real-time detection of traffic signs (i.e., road facilities that transmit guiding, restrictive, warning, or indicating information using words or symbols, which may include main traffic signs and their corresponding smaller-sized auxiliary traffic signs), so as to accurately and reliably control or guide the vehicle's driving based on the traffic indication information contained or represented by the detected traffic signs, and improve the safety of vehicle driving.
[0080] Specifically, a local image acquisition device on the vehicle can be used to obtain a traffic real-scene image in the vehicle's driving direction. Compared with the road image pre-collected and reported by a map collector, the information content of this traffic real-scene image is consistent with the driving road facilities and can include various traffic signs added to the driving road. In this way, the traffic real-scene image is sent to the processor in the local in-vehicle terminal to input a target detection network for traffic sign detection, determine each traffic sign contained in the traffic real-scene image, and then through processing such as semantic or text recognition of the detected traffic signs, obtain the corresponding traffic indication information, so that the vehicle driver or the in-vehicle terminal can timely and accurately know the traffic conditions of the current driving road, and accordingly can timely adjust the vehicle's driving direction, speed, etc., to ensure the safety of vehicle driving and the reliability of navigation.
[0081] However, since the resolution of the traffic real-scene image directly collected by the image acquisition device is usually relatively high, and the traffic signs installed on the road include both main traffic signs with a large area and auxiliary traffic signs with a small area. In this way, in order to ensure reliable detection of the auxiliary traffic signs and avoid the disappearance of the features of the auxiliary traffic signs due to excessive downsampling, the high-resolution traffic real-scene image can be directly input into the target detection network (i.e., the traffic sign detection network) for traffic sign detection. Although it can retain the features of the targets with a small area in the whole image, the computational complexity of this detection process is very large. For terminal devices with poor computing power in local vehicles, such as in-vehicle terminals equipped with an embedded system with a low computing power, etc., it will not be able to support the smooth operation of this detection method, and problems such as freezing and crashing may occur.
[0082] To address the above problems, it is proposed to first perform compression processing such as downsampling on the collected high-resolution traffic real-scene images. After reducing the image resolution, the images are then input into the target detection network for traffic sign detection. However, this method is likely to cause the disappearance of the features of auxiliary traffic signs with relatively small areas, resulting in the failure of target detection. In response to this, the present application further proposes that after performing downsampling processing on the collected traffic real-scene images, first detect the main traffic signs with relatively large areas, and then use the prior knowledge of the detected main traffic signs to determine the potential image regions where the auxiliary traffic signs that match the main traffic signs and are located near the main traffic signs in the traffic real-scene images. At this time, since the traffic signs are the main features of the potential image regions, the present application directly performs auxiliary traffic sign detection on the potential image regions. Compared with directly performing traffic sign detection on the collected high-resolution traffic real-scene images, the computational amount of image processing is greatly reduced, and the detection efficiency and accuracy of the auxiliary traffic signs are improved, thereby ensuring the real-time performance and accuracy of vehicle driving control or guidance.
[0083] Moreover, the traffic sign detection method proposed in the present application occupies relatively low computational resources during operation, can be better applied to terminal devices with relatively low computing capabilities, can be directly implemented by the vehicle-local terminal, and can better meet the detection requirements in different scenarios; and since the in-vehicle terminal does not need to be connected to the network and relies on the map server to implement traffic sign detection, it also avoids data delay or reception failure caused by abnormal network communication, and reduces the real-time performance and accuracy of the traffic sign detection results for vehicle driving automatic control or guidance.
[0084] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts, and in the case of no conflict, the embodiments and the features in the embodiments in the present application can be combined with each other, all belong to the scope protected by the present application.
[0085] It should be understood that the "system", "device", "unit" and / or "module" used in the present application are a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.
[0086] As shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. A method or device may also include other steps or elements. An element defined by the statement "comprising one..." does not exclude the existence of other identical elements in the process, method, commodity, or device that includes the element.
[0087] Among them, in the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; "and / or" herein is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two. The following terms "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0088] In addition, flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the previous or subsequent operations do not necessarily need to be executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0089] Referring to Figure 1 , which is a schematic diagram of the hardware structure of an optional example of the terminal device for implementing the traffic sign detection method provided in this application. In practical applications, the terminal device may include but is not limited to devices with relatively low computing capabilities such as smartphones, tablet computers, desktop computers, netbooks, smart watches, Augmented Reality (AR) devices, Virtual Reality (VR) devices, in-vehicle devices, etc. This application does not limit the product type of the terminal device and can be determined according to the situation. As Figure 1 shown, the terminal device proposed in this embodiment may include but is not limited to: a communication interface 11, a memory 12, and a processor 13, where:
[0090] The number of each of the communication interface 11, the memory 12, and the processor 13 can be at least one, and the communication interface 11, the memory 12, and the processor 13 can be connected to a communication bus, and data interaction can be achieved among them through the communication bus. The specific implementation process can be determined according to actual application requirements, and will not be elaborated in this application.
[0091] The communication interface 11 can be an interface of a communication module applicable to a wireless network or a wired network, such as a communication interface of communication modules such as a GSM module, a WIFI module, a Bluetooth module, a radio frequency module, a 5G / 6G (fifth-generation mobile communication network / sixth-generation mobile communication network) module, etc., and can achieve data interaction with other devices. In this embodiment, it can be used to transmit the traffic real-scene images collected by the image acquisition device, so that the processor 13 obtains the position information of various types of traffic signs in the traffic real-scene images according to the traffic sign detection method proposed in this application, and feeds it back to other devices for output or further processing, etc. The specific implementation can be determined according to actual application requirements, and will not be elaborated in this embodiment of this application.
[0092] It can be understood that the above communication interface 11 can also include interfaces such as a USB interface, a serial / parallel port, etc., for realizing data interaction between internal components of a computer device. Regarding the interface types and quantities included in the communication interface 11, they can be determined according to the device type of the computer device and its application requirements, and will not be elaborated one by one in this application.
[0093] The memory 12 can be used to store a program (which includes multiple computer operation instructions) for implementing the traffic sign detection implementation method proposed in this application; the processor 13 can be used to load and execute the program stored in the memory 12 to implement the steps of the traffic sign detection method proposed in this embodiment of this application. The specific implementation process can refer to but is not limited to the description of the corresponding part of the method embodiment below, and will not be elaborated in this embodiment.
[0094] In this embodiment of this application, the memory 12 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage devices. Specifically, it can be a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM), etc. This application does not limit the type and quantity of the memory 12 included in the computer device, and can be determined according to the situation.
[0095] The processor 13 may include, but is not limited to: a Central Processing Unit (CPU), an application-specific integrated circuit (ASIC), a Digital Signal Processor (DSP), an ASIC, a Field Programmable Gate Array (FPGA), or other programmable logic devices, etc. The type and quantity of devices included in the processor 13 can be determined according to the scenario requirements, and the present application does not limit this.
[0096] In still other embodiments, when the above terminal device is a device such as a smart phone or a vehicle-mounted device listed above, referring to Figure 2 the structural schematic diagram of another optional example of the terminal device shown, the terminal device may further include input devices such as a touch sensing unit for sensing touch events on the inductive touch display panel, a pickup, etc.; at least one output device such as a display, a speaker, a vibration mechanism, a light, etc.; a power supply module; a sensor module; an antenna, etc., which can be determined according to actual needs, and the present application does not elaborate on them one by one here.
[0097] In this way, in the practical application of these still other embodiments, the terminal device efficiently and accurately detects the main traffic signs and their corresponding auxiliary traffic signs in the traffic real-scene images collected locally according to the traffic sign detection method proposed in the present application. In different application scenarios, corresponding application requirement data can be obtained according to the target position information of the detected auxiliary traffic signs in the traffic real-scene images. For example, in a navigation scenario, semantic recognition can be performed on the traffic real-scene image according to the target position information to obtain auxiliary traffic indication information. Then, corresponding traffic reminder information can be generated by combining vehicle driving state information and / or system time, etc., and then output through voice broadcast or pop-up reminder box to remind the driver to adjust the vehicle driving lane, vehicle speed, etc. in time when needed. However, it is not limited to this vehicle navigation application scenario, and the implementation process for other application scenarios is similar, and the present application does not list them one by one.
[0098] Of course, after the main traffic signs and their matching auxiliary traffic signs are detected from the traffic real-scene images collected in real time in the present application, the corresponding map data in the map service device can also be updated according to the detection result. In this way, when accessing the map of this section subsequently, map data including the main traffic signs and their matching auxiliary traffic signs will be obtained, and even the auxiliary traffic reminder information of the auxiliary traffic signs and / or the traffic reminder information obtained therefrom can be displayed by referring to the above method to improve the accuracy and reliability of online map navigation.
[0099] It should be understood that Figure 1The structure of the terminal device shown does not limit the terminal device in the embodiments of the present application. In actual applications, the terminal device may include more or fewer components than Figure 1 shown, or combine certain components, which can be determined according to the device type and functional requirements of the terminal device. The present application does not list them one by one here.
[0100] Referring to Figure 3 , it is a schematic flowchart of an optional example of the traffic sign detection method proposed in the present application. This method can be applied to the above terminal device. The terminal device can be an in-vehicle terminal with relatively low computing power. The present application does not limit the specific product type of the in-vehicle terminal, which can be determined according to the specific application scenario. As Figure 3 shown, the traffic sign detection method proposed in this embodiment may include but is not limited to the following steps:
[0101] Step S11, obtain a traffic real-scene image in the vehicle driving direction;
[0102] Step S12, perform compression processing on the traffic real-scene image to obtain an image to be detected; the resolution of the image to be detected is less than the resolution of the traffic real-scene image;
[0103] Combined with the above description of the technical concept of the present application, in order to ensure reliable and accurate road guidance, the present application uses an image acquisition device in the vehicle, such as a driving recorder, a camera installed on the vehicle windshield with the lens facing the vehicle driving direction, etc., to collect traffic real-scene images in the vehicle driving direction in real time or periodically. The present application does not limit the specific acquisition implementation method of the traffic real-scene image, which can be determined according to the situation.
[0104] After that, the image acquisition device can send the collected traffic real-scene image to the processor of the terminal device for image processing. Specifically, a target detection algorithm can be used to detect various traffic signs therein to achieve road guidance for vehicle driving. Since the area (i.e., scale) of the input image of the target detection network is usually determined, but for different types of image acquisition devices, their working performances are different, and the attribute information such as the image area and / or resolution collected in the same scene may be different. Therefore, the attribute information of the directly collected traffic real-scene image may not meet the input format requirements of the target detection network, and the present application can preprocess the traffic real-scene image according to the input format requirements.
[0105] Meanwhile, considering the limited computing resources that the terminal device can provide, in order to reduce the amount of computation in the image processing, this application proposes to perform compression processing on the directly captured traffic scene image, such as image pixel downsampling, etc., to obtain a to-be-detected image with a specific area and a lower resolution. That is to say, this application can use a lossy compression method to compress the traffic scene image with a higher resolution. The specific compression implementation process will not be elaborated. It can be understood that the image loss caused by this lossy compression method will not affect the detection result of the main traffic sign, that is, within the range of image distortion allowed by this application.
[0106] In some other embodiments, if the resolution of the traffic scene image directly captured by the image acquisition device is relatively low and the image area meets the input image format requirements of the target detection network, it is also possible not to perform compression processing on it and directly use it as the to-be-detected image for subsequent processing. Therefore, in practical applications, it is possible to detect whether the area of the traffic scene image is larger than the preset main traffic sign detection area and whether the resolution is greater than the resolution threshold. If so, perform the above step S12; if the detection results are both negative, the traffic scene image can be determined as the to-be-detected image.
[0107] Step S13: Perform main traffic sign detection on the to-be-detected image, and based on the main traffic sign detection result, obtain a potential image area where there is a first auxiliary traffic sign that matches the detected first main traffic sign;
[0108] For the traffic signs deployed on the traffic road, they usually include the main traffic sign and its matching auxiliary traffic sign. The auxiliary traffic sign is usually located near the main traffic sign. As Figure 4 shown, the auxiliary traffic sign is usually used to explain the time range, road section range, object range, etc. of the function of the main traffic sign to accurately interpret the meaning of the main traffic sign. Its area is usually relatively small. This application does not limit the content of the main traffic sign and its matching auxiliary traffic sign respectively, as well as the relative position relationship between the two, which can be determined according to the situation.
[0109] After the above-mentioned image compression processing, the resolution of the image to be detected is relatively small. For the auxiliary traffic signs that occupy a relatively small area in the entire image to be detected, in order to avoid the disappearance of the features of the auxiliary traffic signs due to excessive downsampling by the target detection network, resulting in detection failure, as described above in the related description of the technical concept of this application, this application will perform target detection on the main traffic signs that occupy a relatively large area in the image to be detected, so as to locate the main traffic signs in the traffic scene image, and then use the adjacent position relationship between the main traffic signs and the auxiliary traffic signs to determine the local image area where the auxiliary traffic signs exist, and form a potential image area for detecting the auxiliary traffic signs. Obviously, compared with the traffic scene image, the potential image area determined by this application contains much less local area, and most of the image features in this potential image area are the image features of traffic signs, which helps to quickly and accurately identify the area where the auxiliary traffic signs are located.
[0110] It should be noted that the above-mentioned main traffic sign detection result may include the regional position of the detected first main traffic sign in the traffic scene image. In order to obtain this main traffic sign detection result, in the embodiments of this application, a main traffic sign detection network may be used to perform main traffic sign detection on the image to be detected, and after locating the area where each main traffic sign (for the convenience of subsequent scheme description, this application will denote the detected main traffic sign as the first main traffic sign) contained in the image to be detected, according to the scaling relationship between the image to be detected and the traffic scene image, locate the area where the first main traffic sign is located in the traffic scene image, and the specific implementation process is not limited in this application.
[0111] It can be seen that the potential image area obtained by this application can be a local image area of the traffic scene image (such as the main traffic sign is located in the relatively middle position of the traffic scene image, etc.), or part of it is a local image area of the traffic scene image, and the other part is a filled image area filled based on this local image area (which is applicable to scenarios such as the first main traffic sign is located at the edge of the traffic scene image). In this case, this potential image area may include this local image area and the filled image area. Therefore, the above-mentioned potential image area may include the corresponding first auxiliary traffic sign, as well as at least part of the first main traffic sign that matches the first auxiliary traffic sign. Specifically, the image content included in the potential image area can be determined according to the actual situation, and this application does not make any restrictions on this. Compared with the entire traffic scene image, the potential image area greatly reduces the included road elements and image features. In this way, performing target detection on this potential image area greatly improves the target detection efficiency and accuracy.
[0112] Step S14: Perform auxiliary traffic sign detection on the potential image region, and based on the auxiliary traffic sign detection result, obtain the target position information of the detected first auxiliary traffic sign mapped on the traffic real-scene image.
[0113] As described above, in the embodiment of the present application, the first main traffic sign included in the traffic real-scene image is located. After obtaining the potential image region of the first auxiliary traffic sign that is located nearby and matches it from the traffic real-scene image, in order to achieve accurate and rapid detection of the first auxiliary traffic sign, the embodiment of the present application uses this potential image region as the input image for subsequent target detection to achieve the detection of targets such as auxiliary traffic signs.
[0114] It can be understood that since the auxiliary traffic sign becomes a large-area target in this potential image region, even if the resolution of the potential image region is low and the target detection network performs excessive downsampling on it, the auxiliary traffic sign can still be successfully sampled, ensuring the reliable implementation of accurate detection of the auxiliary traffic sign on a computer device with low computing power. The present application does not limit the method for implementing the auxiliary traffic sign detection in the potential image region.
[0115] In practical applications, regarding the specific implementation methods of the above-mentioned main traffic sign detection and auxiliary traffic sign detection, in order to further improve the target detection efficiency and accuracy, the present application can adopt technologies included in artificial intelligence (AI) such as machine learning and deep learning, train the corresponding target detection model, and call the corresponding target detection model to implement. Among them, artificial intelligence technology is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0116] Machine learning and deep learning, as the core of artificial intelligence, are the fundamental ways to make a computer intelligent. In practical applications, algorithms such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning can be adopted to implement the learning and training of the corresponding model or network structure to meet specific application requirements. In the embodiment of the present application, algorithms such as convolutional neural network algorithms can be used to train the main traffic sign detection network (i.e., the main traffic sign detection model) / auxiliary traffic sign detection network (i.e., the auxiliary traffic sign detection model) to achieve the detection of the target (i.e., the main traffic sign or the auxiliary traffic sign) in the corresponding input image and determine the region where the corresponding traffic sign is located in the input image. The specific implementation process is not elaborated in the present application.
[0117] It should be understood that if it is detected that the traffic scene image contains multiple first main traffic signs in the above manner, the potential image regions of the first auxiliary traffic signs matching each first main traffic sign can be determined according to the above processing method, that is, the potential image regions of each of the multiple first auxiliary traffic signs are obtained. After that, for each potential image region, target detection can be performed in the above manner to locate the first auxiliary traffic signs actually contained in each potential image region. The detection process is the same, and the embodiments of the present application will not be described in detail one by one.
[0118] As can be seen from the auxiliary traffic sign detection process described above, what is directly detected in the embodiments of the present application is the regional position of the first auxiliary traffic sign, which is actually the regional position of the first auxiliary traffic sign in the potential image region, that is, this regional position refers to the position in the image coordinate system of the potential image region, that is, it is represented by the image pixel position in the potential image region. However, since the potential image region is not the image region directly captured in the traffic scene image, this regional position cannot directly represent the position of the first auxiliary traffic sign in the traffic scene, and further position conversion is required to map it to the traffic scene image before it can be used to implement vehicle road guidance, map hot update, autonomous driving control, etc.
[0119] Therefore, after the position information of the first auxiliary traffic sign in the detected potential image region in the coordinate system of the potential image region is obtained, the target position information of the first auxiliary traffic sign in the traffic scene image can be obtained based on the image position relationship between the potential image region and the traffic scene image, that is, the position information in the coordinate system of the potential image region is converted to the coordinate system of the traffic scene image to obtain the position detection result of the same auxiliary traffic sign in the traffic scene image. It should be noted that the present application does not limit the implementation method of the position coordinate conversion processing of the same region in different image coordinate systems, which can be determined according to the situation.
[0120] Therefore, the above-mentioned target position information obtained by the present application is the regional position of the first auxiliary traffic sign in the traffic scene image, which can represent the coordinate position of the first auxiliary traffic sign on the actual road. In this way, in different application scenarios, the traffic scene image can be processed accordingly, or analyzed in combination with other scene information such as the driving state information of the vehicle (such as the current position of the vehicle, driving speed / acceleration, driving direction, etc.) and system time to obtain reliable and accurate data that meets the requirements of the application scenario, such as a navigation image containing auxiliary traffic signs rendered, traffic prompt information guiding the driving direction / speed of the vehicle, structured data for updating the corresponding map data, etc. The present application does not limit the usage method of the target position information of the auxiliary traffic signs in each frame of traffic scene image detected in different application scenarios, which can be determined according to the situation.
[0121] In some other embodiments, the present application can also utilize computer vision technology to identify traffic signs. Among them, computer vision technology refers to machine vision that uses cameras and computers to replace the human eye to identify and measure targets, etc., and further performs graphic processing to make the computer process into an image that is more suitable for human eye observation or transmission to instrument detection. Therefore, it is usually applied to fields such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality VR, augmented reality AR, simultaneous localization and mapping, etc., as well as biometric recognition application scenarios such as face recognition and fingerprint recognition.
[0122] Based on this, in this embodiment, if it is necessary to further identify the traffic indication information contained or represented by the detected traffic sign, computer vision technology can be used to perform semantic recognition on the traffic sign area image. Subsequently, the obtained traffic indication information, or traffic prompt information generated by combining parameters such as vehicle driving state and system time, etc., can be directly subjected to voice broadcast, or more precisely, the rendering of AR navigation images can be realized, etc., to achieve a more accurate and intelligent navigation plan, or the structured perception results of various traffic signs in the traffic real-scene image can be uploaded to the map service device to update the corresponding map data in real time and improve the accuracy of online map data. The specific implementation process is not elaborated in the present application here.
[0123] In summary, for an in-vehicle terminal with relatively low computing power locally, after obtaining a high-resolution traffic real-scene image in the vehicle driving direction collected locally, it can be compressed to reduce the image resolution, and then the main traffic sign detection can be performed on the obtained low-resolution image to be detected, and the first main traffic sign in the traffic real-scene image can be located. The computational complexity of this target detection process is small, and the computing resources of the terminal device can support the realization of the main traffic sign detection. After that, based on the prior knowledge of the main traffic sign detection result, the potential image area of the first auxiliary traffic sign that exists in the traffic real-scene image and matches the detected first main traffic sign can be obtained, and the auxiliary traffic sign detection can be performed on it, and the auxiliary traffic sign detection result can be obtained quickly and accurately. Based on this, the target position information mapped to the traffic real-scene image can be obtained, so as to ensure the accurate positioning of the auxiliary traffic sign in the traffic real-scene image while greatly reducing the computational complexity of the entire traffic sign detection process to be very small, and further enabling the in-vehicle terminal to realize vehicle control in real time and accurately, improving vehicle driving safety and user driving experience.
[0124] Refer to Figure 5, which is a schematic flowchart of another optional example of the traffic sign detection method proposed in this application. The embodiments of this application can be an optional refined implementation method of the traffic sign detection method described in the above embodiments, but are not limited to the refined implementation method described in this embodiment. For example, Figure 5 As shown, the method may include:
[0125] Step S21, obtain a traffic real-scene image in the vehicle driving direction;
[0126] Step S22, compress and sample the traffic real-scene image according to a preset compression ratio to obtain an image to be detected;
[0127] Regarding the specific implementation processes of step S21 and step S22, reference can be made to the descriptions of the corresponding parts in the above embodiments, and this embodiment will not elaborate.
[0128] Step S23, input the image to be detected into the main traffic sign detection network to obtain the first main traffic signs included in the image to be detected;
[0129] In the embodiments of this application, the main traffic sign detection network is an object detection network for detecting targets such as main traffic signs. Combining the descriptions of the corresponding parts in the above embodiments, it can be pre-trained based on a convolutional neural network (or other machine learning / deep learning networks) for the first sample images containing various main traffic signs until the training constraint conditions are met, such as the number of training times reaches a preset number, the detection accuracy of the trained network reaches an accuracy threshold, or the loss value is less than a loss threshold, etc. The finally trained convolutional neural network is determined as the main traffic sign detection network, that is, the main traffic sign detection model. This application does not elaborate on the specific training implementation process of the main traffic sign detection network.
[0130] In this way, in practical applications, after obtaining the image to be detected with a lower resolution in the above manner, the pre-trained main traffic sign detection network can be called, and the image to be detected obtained this time is input into the main traffic sign detection network. Multiple convolutional layers included in it perform feature extraction, and thus, based on the extracted feature maps, the detection frames of the main traffic signs included in the image to be detected are determined, that is, the position information of the main traffic signs included in the image to be detected is determined.
[0131] Following the above description, since the resolution of the image to be detected is relatively low, the computational amount of feature extraction by the convolutional layer is greatly reduced, and since the size of the main traffic signs is usually relatively large and the area occupied in the entire traffic scene image is relatively large, in this way, even if the traffic scene image is compressed to a lower resolution, the features of the main traffic signs will still be retained, so as to ensure that the main traffic sign detection network can reliably and efficiently identify one or more main traffic signs included in the image to be detected, denoted as the first main traffic signs.
[0132] Step S24: According to a preset compression ratio, perform position correction processing on the first main traffic sign included in the to-be-detected image obtained, and locate the first main traffic sign included in the traffic real-scene image;
[0133] As described above about the technical concept of this application, since the area of the auxiliary traffic sign in the entire to-be-detected image is relatively small, if the target detection network directly performs auxiliary traffic sign detection on the to-be-detected image, it is very easy to cause detection failure due to excessive downsampling. To solve this problem, this application will first perform target detection on the main traffic sign with a larger area in the to-be-detected image to obtain the regional position of the main traffic sign in the to-be-detected image.
[0134] Since the to-be-detected image has been compressed compared to the traffic real-scene image, the number of pixel points included in the to-be-detected image is less than the number of pixel points included in the traffic real-scene image. Therefore, the position representations of the same object in these two images are different. So, in order to more accurately detect the attached traffic sign, when obtaining the local image area where there is an attached traffic sign, it will be obtained from the directly collected traffic real-scene image. In this way, it is necessary to map the regional position of the main traffic sign obtained in the to-be-detected image to the directly collected traffic real-scene image, that is, locate the first main traffic sign in the traffic real-scene image. The specific implementation process is not described in detail in this application.
[0135] Among them, referring to Figure 6 As shown in the schematic diagram of the scenario process, the detection result of the first main traffic sign in the traffic real-scene image can be the position area where the corresponding detection frame is located. Therefore, the position of the first main traffic sign in the traffic real-scene image can be identified by the position where the detection frame is located, and the position of the detection frame can be represented by the coordinates of the diagonal pixel points of the detection frame, but it is not limited to this representation method and can be determined according to the situation.
[0136] Step S25: Taking the first main traffic sign included in the traffic real-scene image as a reference object, determine the local image area where there is a first auxiliary traffic sign that matches the first main traffic sign;
[0137] Continuing the above description, after determining the area where the first main traffic sign in the traffic real-scene image is located, as shown above Figure 4 Since the auxiliary traffic sign is usually located near the main traffic sign, this application can determine the local image area where the first auxiliary traffic sign in the traffic real-scene image is located with reference to the position of the first main traffic sign in the traffic real-scene image.
[0138] In a possible implementation manner, referring to Figure 6The schematic diagram of the scene processing flow shown. In the present application, in a traffic real-scene image, a preset distance can be diffused around the first main traffic sign as the center, and the formed closed area can be determined as the local image area where the first auxiliary traffic sign matching the first main traffic sign exists. However, it is not limited to this implementation method of determining the local image area, and the present application does not limit the specific value of the preset distance, which can be determined according to but not limited to the relative position relationship between the main traffic sign and the auxiliary traffic sign, as well as factors such as their respective sizes. Moreover, the local image area can be a circular area, a square area, etc. The present application does not limit the shape of the local image area and can be determined according to the situation.
[0139] In another possible implementation method, during the diffusion process around the first main traffic sign as the center, the diffusion distances in different diffusion directions can be different, not limited to the diffusion method of the same distance (such as the above-mentioned preset distance) described in the above embodiments. The local image area determined according to this different diffusion method may be a rectangle, an ellipse, an irregular figure, etc., and the present application will not elaborate on them one by one here.
[0140] It should be noted that during the diffusion process around the first main traffic sign as the center, as Figure 7 shown in the traffic real-scene image in the upper attached figure of the scene flow schematic diagram. If the first main traffic sign is located near a certain boundary of the traffic real-scene image, then, when performing area diffusion in the corresponding first diffusion direction, the first diffusion direction will first reach the boundary of the traffic real-scene image, that is, it is detected that the first boundary distance between the first main traffic sign and the first image boundary is less than the preset distance; the first image boundary refers to the boundary of the traffic real-scene image in the first diffusion direction. At this time, in order to continue the diffusion and the diffused area has image features, when diffusing from the first main traffic sign to the first image boundary in the first diffusion direction, the present application can adopt the first image filling method to continue the image filling diffusion in the first diffusion direction until the diffusion distance in the first diffusion direction reaches the preset distance or the preset diffusion distance corresponding to the first diffusion direction, etc. The present application does not elaborate on the image filling implementation process included in the first image filling method.
[0141] It can be seen that during the process of determining the image area where the first auxiliary traffic sign exists in this embodiment, when the local image area determined from the traffic real-scene image according to the above method still does not meet the diffusion requirements, image filling diffusion will continue on the basis of this local image area, and the specific implementation process includes but is not limited to the implementation methods described above.
[0142] In yet another possible implementation, when actually installing the main traffic signs and auxiliary traffic signs, it is usually carried out according to specific rules, so that the positional deployment relationship between the main traffic signs of the same sign type and their auxiliary traffic signs is basically fixed, such as the up-down relationship or the left-right relationship, etc. Therefore, during the process of determining the local image region, the present application can also obtain the first positional deployment relationship corresponding to the main traffic signs of the sign type according to the sign type of the first main traffic sign; wherein, the first positional deployment relationship refers to the relative positional relationship between the main traffic signs of this sign type and the auxiliary traffic signs that match the main traffic signs in the traffic road facilities.
[0143] After that, in the directly captured traffic real-scene image, according to the first positional deployment relationship, determine the local image region (or the image region to be processed composed of the local image region and the above-mentioned image filling region) where there is a first auxiliary traffic sign that matches the first main traffic sign, so as to improve the determination efficiency of the local image region (or the image region to be processed), and ensure that the image features of the first auxiliary traffic sign included in the local image region are as complete and detailed as possible, which helps to improve the accuracy and detection efficiency of subsequent auxiliary traffic sign detection.
[0144] Step S26, form a potential image region from the determined local image region;
[0145] Continuing from the above description, the potential image region to be obtained by the present application can be a local image region in the traffic real-scene image, or may be composed of a part that is a local image region in the traffic real-scene image and another part that is an image filling region obtained by filling the image based on the local image region, which can be determined according to the specific situation. It should be understood that no matter which acquisition method, the potential image region is much smaller than the directly captured traffic real-scene image in terms of image area, that is, it reduces the types and quantities of image features included, which is beneficial to improving the detection efficiency and accuracy of auxiliary traffic signs when performing target detection on the potential image region subsequently.
[0146] In practical applications, the change of the current lens configuration parameters of the image acquisition device and its distance change from the traffic signs will affect the area size of each traffic sign in the captured traffic real-scene image. Therefore, the area of the local image region obtained according to the above processing method will be affected by the area size of the main traffic signs. If the area of the first main traffic sign in the detected traffic real-scene image is large, the area of the local image region obtained accordingly will also be relatively large, which does not meet the input image area requirements of the target detection network; on the contrary, if the area of the obtained local image region is too small, it also does not meet the input image area requirements of the target detection network.
[0147] Therefore, after obtaining the local image region, preprocessing can be performed on the local image region according to the input image format requirements of the auxiliary traffic sign detection network. For example, the specific area required for the input image of the auxiliary traffic sign detection network can be recorded as the preset auxiliary traffic sign detection area. The area of the determined local image region can be compared with the preset auxiliary traffic sign detection area. If the area of the local image region is smaller than the preset auxiliary traffic sign detection area, the local image region is enlarged to obtain a potential image region with the preset auxiliary traffic sign detection area; conversely, if the area of the local image region is larger than the preset auxiliary traffic sign detection area, the local image region is compressed to obtain a potential image region with the preset auxiliary traffic sign detection area.
[0148] It can be seen that after preprocessing the local image, the resolution of the local image region will be reduced, so that the resolution of the obtained potential image region is smaller than that of the traffic real-scene image, thereby reducing the computational complexity of subsequent target detection. Of course, if you want to further reduce the computational complexity of the auxiliary traffic sign detection, the present application can further downsample the local image region to reduce its image resolution, and the specific implementation can be determined according to the situation.
[0149] In some other embodiments, in the scenario where image filling processing needs to be performed on the local image region, the processed image region obtained after image filling processing can be processed according to the above scaling processing method to obtain a potential image region that meets the input image requirements of the auxiliary traffic sign detection network. The specific implementation process is similar, and the present application will not elaborate here.
[0150] It should be noted that the resolution, area, and other attribute parameters of the potential image region and the above-mentioned image to be detected can be the same or different, which can be determined according to the situation, and the present application does not limit this.
[0151] Step S27: Input the potential image region into the auxiliary traffic sign detection network to obtain the first auxiliary traffic sign included in the potential image region and the regional position information of the first auxiliary traffic sign in the potential image region;
[0152] In the embodiment of the present application, the auxiliary traffic sign detection network is a target detection network targeting auxiliary traffic signs, which can be trained based on a convolutional neural network for sample potential images, but is not limited to this acquisition method. Since its construction and implementation process is similar to that of the above-mentioned main traffic sign detection network, the description of the acquisition process of the main traffic sign network above can be referred to, and this embodiment will not be elaborated here.
[0153] Step S28: Obtain the potential regional position information of the potential image region in the traffic real-scene image;
[0154] Step S29: According to the determination method of the potential image region and the potential region position information, perform coordinate conversion processing on the region position information to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0155] As described in the corresponding part of the above embodiments, after obtaining the region position information of the first auxiliary traffic sign in the directly located potential image region, it is necessary to further determine the target position information of the first auxiliary traffic sign mapped to the directly captured traffic real-scene image, that is, convert the region position information of the first auxiliary traffic sign from the image coordinate system represented by the potential image region to the image coordinate system represented by the traffic real-scene image, and obtain the target position information represented by the pixel point position of the traffic real-scene image. This application details the implementation process of how to perform position information conversion processing between two image coordinate systems for the same object.
[0156] In some embodiments, in combination with the process of obtaining the potential image region described in the above embodiments, the process of obtaining the above target position information may include but is not limited to the following steps:
[0157] Determine the image filling data of the traffic real-scene image and the scaling ratio of the local image region with the first auxiliary traffic sign obtained from the traffic real-scene image during the process of obtaining the potential image region; wherein, the image filling data may include the filling distances in different diffusion directions; as analyzed above, the potential image region is obtained by compressing or magnifying the local image region, and the scaling ratio may include the compression ratio or the magnification ratio, which can be determined according to the specific situation. Then, the obtained region position information can be restored by using the filling distances in different diffusion directions, the scaling ratio, and the determined potential region position data to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0158] Exemplarily, assume that the potential region position data of the potential image region is (x 1c , y 1c , x 2c , y 2c ), that is, the position coordinates of the two diagonal pixel points of the potential image region are (x 1c , y 1c ) and (x 2c , y 2c) to represent the potential area position data of the potential image area, and the width and height of the potential image area are compressed or enlarged by S times relative to the directly obtained local image area, that is, the above scaling ratio is 1 / S. If there is an image filling process during the acquisition of the local image area, the filling distances corresponding to the top, bottom, left, and right of the first main traffic sign are respectively recorded as padup, paddown, padleft, and padright.
[0159] It can be understood that if there is no image filling in a certain diffusion direction, then the filling distance in this direction is 0. As shown above Figure 7 In the process of diffusing to the right with the first main traffic sign as the center, after reaching the boundary of the traffic real scene image, it is necessary to fill the subsequent diffusion area. It can be seen that the filling distance corresponding to the right side is the vertical distance between the right side of the traffic real scene image and the right side of the local image area, and the specific value is not limited. By analogy, for other filling scenarios, multiple diffusion directions may need to fill some diffusion areas, thus obtaining multiple filling distances. The acquisition process is similar and will not be elaborated in this application.
[0160] Based on this, if the area position information of the first auxiliary traffic sign in the potential image area is detected as (x1, y1, x2, y2) in the above manner, then when mapping it to the directly captured traffic real scene image, the coordinate values in the target position information (X1, Y1, X2, Y2) of the area where the first auxiliary traffic sign is located can be calculated according to the following formula:
[0161] X1 = (x1 - padleft) / S + x 1c ; Y1 = (y1 - padup) / S + y 1c ;
[0162] X2 = (x2 - padright) / S + x 2c ; Y2 = (y2 - paddown) / S + y 2c ;
[0163] It should be noted that the implementation process of mapping the position of the first auxiliary traffic sign in the potential image area to the traffic real scene image includes but is not limited to the calculation method described above.
[0164] To sum up, during the vehicle driving process, referring to Figure 8 the scene process schematic diagram shown, the traffic real scene image in the driving direction can be collected by the image acquisition device in the vehicle, such as Figure 8As shown in the first picture on the left of the first row, the traffic scene image includes the road in the driving direction and the facilities on both sides of the road, such as road elements like traffic signs, transmission lines, street lights, cameras, etc. Since the resolution of this traffic scene image is often relatively high, if target detection is directly performed, it will require extremely large computing resources and easily cause abnormalities such as the in-vehicle terminal with low computing power crashing or freezing, and it cannot operate normally. Therefore, in this application, the traffic scene image will first be subjected to downsampling compression processing to obtain a to-be-detected image with a low resolution and a preset size.
[0165] After that, first perform target detection on the main traffic signs in the to-be-detected image where the traffic sign features are relatively obvious and relatively detailed and complete. After locating the position of the detected first main traffic sign in the originally collected traffic scene image, spread around the detected first main traffic sign to form a potential image area where there are first auxiliary traffic signs near it, so as to realize the target detection of the auxiliary traffic signs.
[0166] Among them, before detection, in order to further improve the detection efficiency and meet the format requirements of the target detection network for the input image, the obtained potential image area can be scaled and then input into the pre-trained auxiliary traffic sign detection network for target detection to determine the regional position of the first auxiliary traffic sign in the potential image area. Then, based on the image coordinate conversion relationship between the potential image area and the traffic scene image, map the auxiliary traffic sign detection result to the traffic scene image, so as to quickly and accurately realize the positioning detection of the auxiliary traffic signs in the traffic scene image locally, without relying on the communication network, and can also meet the application requirements for the positioning detection of the auxiliary traffic signs in the corresponding application scenarios.
[0167] In some other embodiments proposed in this application, after completing the detection of the main traffic signs in the to-be-detected image according to the processing method described in the above embodiments and directly obtaining the first main traffic sign in the to-be-detected image, this application can also directly determine, in this to-be-detected image, a potential image area where there is a first auxiliary traffic sign that matches it with the detected first main traffic sign detection frame as a reference object. At this time, since the to-be-detected image is a low-resolution image, therefore, in this embodiment, the local image area storing the first auxiliary traffic sign intercepted from the to-be-detected image can be directly used as the potential image area without the need for further downsampling to reduce the image resolution. After that, the implementation process of detecting the auxiliary traffic signs in the potential image area is similar, and this embodiment will not be elaborated in detail.
[0168] It should be noted that in the process of obtaining the target position information of the first auxiliary traffic image detected and mapped to the directly captured traffic real-scene image, this embodiment needs to consider the compression ratio in the process of obtaining the image to be detected by compressing the traffic real-scene image. That is, according to this compression ratio, the potential area position information of the potential image area in the image to be detected, and data such as the filling distance in different diffusion directions for obtaining the potential image area, coordinate transformation processing is performed on the area position information of the first auxiliary traffic sign in the detected potential image area to obtain its target position information in the traffic real-scene image. The specific transformation process is not described in detail in this embodiment.
[0169] It should be understood that in practical applications, in order to improve the detection accuracy of auxiliary traffic signs, this application preferably selects the method described above of extracting the local image area where the first auxiliary traffic sign exists from the directly captured traffic real-scene image to construct the potential image area.
[0170] Based on the traffic sign detection methods described in the above embodiments, the implementation process of this method in corresponding application scenarios will be described below by taking different application scenarios such as navigation scenarios, map hot updates, and autonomous driving as examples.
[0171] Exemplarily, in the AR navigation scenario, it combines the navigation positioning technology with the AR technology to provide a new navigation form. Among them, the AR technology is a technology that increases the user's perception of the real world through virtual information provided by a computer system. That is, the AR technology can apply virtual information to the real world, realizing the superposition of virtual objects, virtual scenes or system prompt information generated by a computer onto the real scene, thereby realizing the enhancement of reality. In other words, the goal of this technology is to overlay the virtual world on the real world on the screen and interact. In addition, since the real scene and virtual information are superposed on the same screen in real time, after being perceived by human senses, it can achieve a sensory experience beyond reality.
[0172] It can be seen that running the AR technology in the field of map navigation realizes AR navigation in addition to traditional two-dimensional map navigation, and can superimpose and display virtual navigation prompt information in the real scene image, so that users can obtain a sensory experience beyond reality. It should be noted that this AR navigation technology can be applied not only in the driving scenario but also in other scenarios with map navigation requirements such as walking. This application does not specifically limit this, and this embodiment only takes the AR navigation in the driving scenario as an example for illustration.
[0173] Combined with the above description of the technical concept of the present application, the present application hopes to achieve AR navigation without relying on the network and the server, but instead, the local terminal device accurately detects various types of traffic signs on the driving road in real time to achieve accurate and efficient rendering of road images, meeting the real-time and accuracy requirements of AR navigation.
[0174] Specifically, when the user needs to turn on AR navigation during the vehicle driving process, the user can input the destination into the terminal device for navigation operations, generating a corresponding AR navigation request to request the target navigation data for navigating to the destination. Based on this AR navigation request, the local image acquisition device can be triggered to start. The image acquisition device captures images of the road in the driving direction of the vehicle to obtain real-time traffic scene images. Then, as described in the traffic sign detection method in the above embodiment, combined with Figure 8 as shown, the downsampling and compression processing is performed on any frame of the real-time captured traffic scene image, and the obtained detected image with a preset size and lower resolution is input into the main traffic sign detection network, and the detection box where the main traffic sign (such as Figure 8 the 40th as shown) in this frame of traffic scene image is output. In order to further detect the auxiliary traffic signs located nearby (that is, the auxiliary information used to explain the meaning of the main traffic sign, such as speed limit, after 19:00, 200 meters ahead, etc., depending on the situation), taking the detection box where the main traffic sign is detected as a reference, a local area is divided around it as the potential image area where there may be auxiliary traffic signs, and then the auxiliary traffic signs in this potential image area are directly detected. Compared with directly detecting the auxiliary traffic signs from the entire traffic scene image, the image processing workload is greatly reduced, the detection efficiency and accuracy are improved, and it can be better applied to in-vehicle terminals with poor computing power.
[0175] Among them, during the detection process of the auxiliary traffic signs, in order to further reduce the calculation amount, the local image area divided from the traffic scene image can be first downsampled to obtain a potential image area with a specific size and lower resolution, and then input into the auxiliary traffic sign detection network for detecting the auxiliary traffic signs. The position of the auxiliary traffic signs in the potential image area can be quickly and accurately determined, and then combined with the compression / enlargement, image filling and other processing methods performed on the potential image area obtained before, the position information of the detected auxiliary traffic signs is mapped to the traffic scene image, that is, the present application accurately detects the main traffic signs and their matching auxiliary traffic signs in the traffic scene image of the current road. The entire detection process is realized on the local in-vehicle terminal without relying on a server with strong computing power through the network, avoiding various problems caused by poor network quality; moreover, the present application detects traffic signs on the real-time captured traffic scene image of the driving road, which is more accurate than the road scene images reported by road collectors to the server.
[0176] After that, the in-vehicle terminal can render the traffic real-scene image of the current frame based on the obtained target position information of the auxiliary traffic signs in the traffic real-scene image, ensuring that the rendered navigation image is consistent with the road scene in the vehicle driving direction, including various road facilities and auxiliary facilities on the road in detail, such as speed limit signs and their attached sign contents, turning signs, traffic lights, speed cameras, cameras, lane lines, road edge lines, etc. In this way, whether the navigation image rendered is directly output by the in-vehicle terminal display, or the in-vehicle terminal projects the navigation image onto an object such as a windshield for display, or other output methods, it can help the driver see various facilities on the road ahead more clearly and accurately, so as to adjust the driving route or driving speed, etc., to drive safely in accordance with traffic rules.
[0177] Among them, in order to improve the reminder reliability and timeliness of the auxiliary traffic indication information of the auxiliary traffic signs, the present application can pop up a traffic reminder window in the navigation image, and the traffic reminder window displays the traffic reminder information for the first auxiliary traffic sign, ensuring that the driver can quickly see the traffic reminder information; it should be noted that the traffic reminder information can be determined based on the auxiliary traffic indication information of the first auxiliary traffic sign, and parameters such as the driving state information and / or system time of the vehicle, and the specific content can be determined according to the situation, and the present application does not limit this.
[0178] Exemplarily, if a traffic sign restricting traffic from 12:00 to 16:30 for the next 500 meters is temporarily set for a certain section ahead of the vehicle, the map data provided by the map server does not include the content of this traffic sign. In this way, in traditional AR navigation applications, the navigation route provided by the map server will be inaccurate and cannot timely remind the driver to detour, which is very inconvenient. However, by using the traffic sign detection method proposed in the present application, the vehicle's local image acquisition device can collect the traffic real-scene image of the driving road in real time, and timely obtain the traffic sign with the above content set temporarily. Specifically, according to the above traffic sign detection method, accurately locate the main traffic sign in the temporarily set traffic signs in the traffic real-scene image, and can also accurately locate the matching auxiliary traffic signs. Then, based on the detection results, perform semantic recognition on each detected traffic sign, and the traffic indication information contained in the traffic sign can be obtained, such as restricting traffic from 12:00 to 16:30 for the next 500 meters.
[0179] After that, the in-vehicle terminal can obtain the current system time and the relative distance between the current vehicle and the traffic sign. Then, by combining the traffic indication information obtained from the above detection, the distance between the restricted section and the vehicle can be calculated. At the same time, it is determined whether the current system time is within the restricted time period. If it is within the restricted time period, traffic reminder information indicating that the section ahead of the current time is restricted can be output. For example, the restriction starts 700 meters ahead at the current time. According to needs, information such as the time when the restriction is lifted can also be output. This application does not limit the content of this traffic reminder information. For this traffic reminder information, as analyzed above, it can be displayed by a traffic reminder window popped up on the display screen, or output in the form of voice broadcast. This application does not limit its output method.
[0180] It can be understood that for the detected auxiliary traffic signs containing other content, the content of the traffic reminder information obtained accordingly will also change. This application does not list them one by one here and can be determined according to the situation.
[0181] In some other embodiments, such as the AR navigation scenario described above, this application can also perform image rendering based on the positions of the main traffic sign and the auxiliary traffic sign in the detected traffic real-scene image. It can ensure that the traffic signs with the above content can be accurately displayed in the rendered and displayed navigation image. According to needs, the display state of the traffic sign can also be adjusted to highlight the content of the traffic sign. For example, the auxiliary traffic sign can be enlarged for display to ensure that the driver can see in time the content of "restricted from 500 meters ahead at 12:00 to 16:30" contained in the traffic sign and adjust the driving route in time. It should be noted that the detection and output methods for traffic signs with other content are similar, and this application does not elaborate on them one by one here.
[0182] Thus, this application uses the local in-vehicle terminal to implement AR navigation, directly collects the traffic real-scene images of the driving road, performs secondary detection of the target object based on the partial image determined by the main traffic sign contained therein, and can quickly and accurately locate the auxiliary traffic signs in the traffic real-scene image with fewer resources. In this way, image rendering is performed based on the located auxiliary traffic signs, which can ensure that the rendered navigation image can accurately and completely display the auxiliary traffic signs, improving the reliability and accuracy of AR navigation.
[0183] In still other embodiments, in the above navigation scenario, the present application can also obtain the relative distance between the vehicle and the corresponding first auxiliary traffic sign according to the target position information, and can also obtain the main traffic indication information of the first main traffic sign included in the traffic real-scene image, as well as the auxiliary traffic indication information of the first auxiliary traffic sign matching the first main traffic sign. Then, using the relative distance, the main traffic indication information, and the auxiliary traffic indication information, a road guidance message in the vehicle driving direction is generated. After that, the road guidance message can be broadcast by voice, and of course, other output methods can also be used to output the road guidance message. The present application does not limit the output method and presentation form of the road guidance message.
[0184] Exemplarily, if there is a traffic sign (i.e., a traffic board) 100 meters ahead of the vehicle, such as Figure 4 the third traffic sign on the left. According to the method for obtaining the road guidance message described in this embodiment, the obtained road guidance message can be that the traffic sign 100 meters ahead indicates a speed limit of 40 from 8:00 to 18:00. After receiving this road guidance message, the driver can timely adjust the vehicle driving speed so as not to violate the regulations when entering the speed limit section. Compared with being able to see the traffic sign only when reaching the speed limit section, which may cause the driver to react untimely and exceed the speed limit or have accidents such as collisions, the driving safety is improved. The indication processing process for other types of traffic signs is similar, and the present application does not list them one by one.
[0185] In still other alternative examples, such as in the map hot update application scenario, the present application can obtain the auxiliary traffic indication information of the corresponding first auxiliary traffic sign according to the obtained target position information. Then, using the target position information and the auxiliary traffic indication information, a structured perception result of the corresponding first auxiliary traffic sign in the traffic real-scene image is obtained. Regarding the data format and content requirements of the structured perception result, they can be determined according to the map update data format requirements, and the present application does not limit this. After that, the terminal device can report the structured perception result to the map service device, and the map service device uses the structured perception result to update the corresponding map data, thereby improving the accuracy of the corresponding road map data on the map. In this way, in subsequent scenarios where other users request online map navigation, etc., the requested map data can include traffic signs on each road, improving the accuracy of the navigation planning route. Among them, for various types of traffic signs included in the road, the output methods can refer to but are not limited to the voice broadcast, pop-up traffic prompt box, local magnification display, etc. listed above to ensure that the driver can clearly and timely learn the traffic indication information.
[0186] In some other alternative examples, such as in the scenario of autonomous driving, after obtaining the target position information of the auxiliary traffic signs in the current frame of traffic real - scene image in the above - mentioned manner and identifying the auxiliary traffic indication information of the auxiliary traffic signs accordingly, the processor can combine the auxiliary traffic indication information to generate an autonomous driving control instruction for the vehicle, such as control strategies for adjusting the vehicle driving direction, driving speed / acceleration, etc., and send it to the vehicle maneuvering components for execution to adjust the vehicle driving state and ensure the driving safety and reliability of the vehicle.
[0187] In practical applications, in the traffic sign detection method proposed in this application, the obtained traffic real - scene image containing the main traffic sign detection frame and the auxiliary traffic sign detection frame, and even the map data after map update based on the located and detected auxiliary traffic signs and their target position information, etc. can all be saved in the blockchain to improve data storage security, and at the same time, it is also convenient for other users to access and query the latest map data at any time to meet the application requirements.
[0188] Refer to Figure 9 , which is a schematic structural diagram of an alternative example of the traffic sign detection device proposed in this application. This device can be applicable to terminal devices, such as AR glasses, AR helmets, tablet computers or other in - vehicle devices with relatively poor computing capabilities, depending on the situation. This application does not limit the product type of this terminal device. As Figure 9 shown, the device may include:
[0189] A traffic real - scene image acquisition module 21, configured to acquire a traffic real - scene image in the vehicle driving direction;
[0190] A to - be - detected image obtaining module 22, configured to perform compression processing on the traffic real - scene image to obtain a to - be - detected image; the resolution of the to - be - detected image is less than the resolution of the traffic real - scene image;
[0191] A potential image area obtaining module 23, configured to perform main traffic sign detection on the to - be - processed traffic image, and obtain a potential image area where a first auxiliary traffic sign that matches the detected first main traffic sign exists according to the main traffic sign detection result; the potential image area includes a local image area of the traffic real - scene image;
[0192] A target position information acquisition module 24, configured to perform auxiliary traffic sign detection on the potential image area, and obtain the target position information of the detected first auxiliary traffic sign mapped in the traffic real - scene image according to the auxiliary traffic sign detection result.
[0193] In some embodiments, as Figure 10 shown, the above - mentioned potential image area obtaining module 23 may include:
[0194] The main traffic sign detection unit 231 is configured to input the image to be detected into the main traffic sign detection network, and obtain the first main traffic sign included in the traffic scene image;
[0195] The local image area determination unit 232 is configured to use the first main traffic sign included in the traffic scene image as a reference object, and determine the local image area of the first auxiliary traffic sign that matches the first main traffic sign;
[0196] In some embodiments, the local image area determination unit 232 may include:
[0197] The first diffusion processing unit is configured to, in the traffic scene image, diffuse a preset distance around the obtained first main traffic sign as the center, and obtain the local image area of the first auxiliary traffic sign that matches the first main traffic sign.
[0198] In another possible implementation manner, if the first main traffic sign is close to the edge position of the traffic scene image, the local image area determination unit 232 may further include:
[0199] The first detection unit is configured to, in the process of diffusing a preset distance around the obtained first main traffic sign as the center, detect that the first boundary distance between the first main traffic sign and the first image boundary is less than the preset distance; the first image boundary refers to the boundary of the traffic scene image in the first diffusion direction;
[0200] The first filling unit is configured to, in the first diffusion direction, diffuse from the first main traffic sign to the first image boundary, and continue to perform image filling diffusion according to the first image filling method until the diffusion distance in the first diffusion direction reaches the preset distance. In some other embodiments, the above-mentioned local image area determination unit 232 may also include:
[0201] The first position deployment relationship acquisition unit is configured to obtain the first position deployment relationship corresponding to the main traffic sign of the sign type according to the sign type of the first main traffic sign included in the traffic scene image; wherein, the first position deployment relationship refers to the relative position relationship between the main traffic sign of the sign type and the auxiliary traffic sign that matches the main traffic sign in the traffic road facilities;
[0202] The second diffusion processing unit is configured to determine the local image area of the first auxiliary traffic sign that matches the first main traffic sign from the traffic scene image according to the first position deployment relationship.
[0203] The potential image area composition unit 233 is configured to compose a potential image area from the determined local image area.
[0204] As described above, the potential image region forming unit 233 may include:
[0205] A first forming unit for directly forming a potential image region from the determined local image region;
[0206] A second forming unit for forming a potential image region with the first auxiliary traffic sign from the local image region and the image filling region in the obtained traffic real-scene image.
[0207] In a possible implementation manner, when forming a potential image region from a local image region, the potential image region forming unit 233 may include:
[0208] A comparison unit for comparing the area of the determined local image region with a preset auxiliary traffic sign detection area;
[0209] A first obtaining unit for, when the area of the local image region is smaller than the preset auxiliary traffic sign detection area, performing a magnification process on the local image region to obtain a potential image region having the preset auxiliary traffic sign detection area;
[0210] A second obtaining unit for, when the area of the local image region is larger than the preset auxiliary traffic sign detection area, performing a compression process on the local image region to obtain a potential image region having the preset auxiliary traffic sign detection area;
[0211] Wherein, the resolution of the potential image region is smaller than the resolution of the traffic real-scene image.
[0212] Similarly, when forming a potential image region with the first auxiliary traffic sign from the local image region and the image filling region in the obtained traffic real-scene image, the area of the to-be-processed image region obtained by summing the local image region and the image filling region may be compared with the preset auxiliary traffic sign detection area in the above manner, and thus, according to the comparison result, the to-be-processed image region is magnified or compressed to obtain a potential image region.
[0213] Based on the descriptions of the above embodiments, in some other embodiments, as Figure 10 shown, the above-mentioned target position information obtaining module 24 may include:
[0214] An auxiliary traffic sign detection unit 241 for inputting the potential image region into an auxiliary traffic sign detection network to obtain the first auxiliary traffic sign included in the potential image region and the region position information of the first auxiliary traffic sign in the potential image region;
[0215] It can be understood that in the case where the number of the above-detected first main traffic signs is multiple, the above potential image area obtaining module 23 can specifically be used to obtain, according to the detected multiple first main traffic signs, the potential image areas where the respective first auxiliary traffic signs matched with the multiple first main traffic signs exist; wherein, the potential image area includes the corresponding first auxiliary traffic sign and at least part of the first main traffic sign matched with the first auxiliary traffic sign.
[0216] Correspondingly, the above auxiliary traffic sign detection unit 241 can perform auxiliary traffic sign detection on each of the obtained potential image areas.
[0217] The potential area position information obtaining unit 242 is configured to obtain the potential area position information of the potential image area in the traffic real-scene image.
[0218] The coordinate conversion processing unit 243 is configured to perform coordinate conversion processing on the area position information according to the obtaining manner of the potential image area and the potential area position information, so as to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0219] Based on the description of the obtaining process of the potential image area in the above embodiment, in a possible implementation manner, the above coordinate conversion processing unit 243 may include:
[0220] The first data determination unit is configured to determine the image filling data of the traffic real-scene image and the scaling ratio of the local image area where the first auxiliary traffic sign exists obtained from the traffic real-scene image during the process of obtaining the potential image area.
[0221] Wherein, the image filling data includes the filling distances in different diffusion directions; the potential image area is obtained by compressing or enlarging the local image area.
[0222] The first position restoration processing unit is further configured to use the filling distances in different diffusion directions, the scaling ratio and the potential area position data to perform restoration processing on the obtained area position information, so as to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
[0223] It can be understood that if image filling processing is not required during the acquisition of the potential image area, then the above-mentioned first data determination unit may not need to acquire the image filling data for the traffic real-scene image. In this way, the first position restoration processing unit can specifically use the scaling ratio and the potential area position data to perform restoration processing on the obtained area position information, so as to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image. The implementation process is not elaborated in this application.
[0224] Based on the technical solutions described in the above embodiments, in the navigation scenario, the above-mentioned device may further include:
[0225] An image rendering module, configured to render the traffic real-scene image according to the target position information;
[0226] An image output module, configured to output the rendered navigation image, and display the detected first main traffic sign and the first auxiliary traffic sign matching the first main traffic sign in the navigation image.
[0227] In still other embodiments, such as in the above navigation scenario, the above-mentioned device may also include:
[0228] A distance acquisition module, configured to obtain the relative distance between the vehicle and the corresponding first affiliated traffic sign according to the target position information;
[0229] A traffic indication information acquisition module, configured to acquire the main traffic indication information of the first main traffic sign included in the traffic real-scene image and the auxiliary traffic indication information of the first auxiliary traffic sign matching the first main traffic sign;
[0230] A road guidance message generation module, configured to generate a road guidance message in the driving direction of the vehicle by using the relative distance, the main traffic indication information, and the auxiliary traffic indication information;
[0231] A road guidance message broadcast module, configured to perform voice broadcast on the road guidance message.
[0232] Optionally, in the map hot update scenario, the above-mentioned device may also include:
[0233] An auxiliary traffic indication information acquisition module, configured to obtain the auxiliary traffic indication information of the corresponding first auxiliary traffic sign according to the target position information;
[0234] A structured perception result acquisition module, configured to obtain the structured perception result of the corresponding first auxiliary traffic sign in the traffic real-scene image by using the target position information and the auxiliary traffic indication information;
[0235] The structured perception result reporting module is used to report the structured perception result to the map service device, and the map service device uses the structured perception result to update the corresponding map data.
[0236] It should be noted that various modules, units, etc. in the above device embodiments can be stored in the memory as program modules, and the processor executes the above program modules stored in the memory to implement corresponding functions. For the functions implemented by each program module and its combination, and the achieved technical effects, reference can be made to the description of the corresponding part of the above method embodiments, which will not be elaborated in this embodiment.
[0237] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the above traffic sign detection method is implemented. The implementation process of the traffic sign detection method can be referred to the description of the above method embodiments.
[0238] The present application also proposes a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners in the above aspects of the traffic sign detection method and the traffic sign detection device. The specific implementation process can be referred to the description of the above corresponding embodiments and will not be elaborated.
[0239] Refer to Figure 11 As shown in the figure, it is a schematic diagram of the system architecture of an optional application environment applicable to the traffic sign detection method and device proposed in the present application. In this application environment (such as vehicle networking, etc.), its system architecture may include multiple terminal devices 31, multiple image acquisition devices 32, and a map service device 33; among them:
[0240] Regarding the composition structure and functions of the terminal device 31, reference can be made to the description of the above terminal device embodiments, which will not be elaborated in this application. It should be noted that the terminal device may include a device with relatively low computing power installed in the vehicle, which executes the traffic sign detection method described in the above method embodiments of the present application to meet the positioning detection requirements of auxiliary traffic signs in different application scenarios, and the implementation process will not be elaborated.
[0241] The image acquisition device 32 can be used to acquire traffic real-scene images, specifically including driving records located in the vehicle, independent cameras installed on the front windshield, or cameras configured in the terminal device, etc. The present application does not limit the product type of the image acquisition device 32.
[0242] The map service device 33 can be a service device that provides map navigation services. It can be an independent physical server, a service cluster integrated by multiple physical servers, or a cloud server with certain cloud computing capabilities. It can establish a communication connection with the terminal device on the vehicle through a wired network or a wireless network to receive the road guidance data obtained and sent by the terminal device. Of course, in some embodiments, the map service device 33 can also respond to the navigation request initiated by the terminal device and provide the latest map navigation data, etc. The specific implementation process is not described in detail in this application.
[0243] Based on this, a map application program can be installed in the above terminal device to start the map application program, access the above map server, and view map data. Such as various taxi application programs, dedicated map application programs, etc., which can be determined according to actual application requirements.
[0244] It should be understood that Figure 11 the system structure shown does not constitute a limitation to the system in the embodiments of this application. In actual applications, the system architecture can include more or fewer components than Figure 11 shown, or combine some components, such as data storage devices, etc., which can be determined according to the specific requirements of the application scenario. This application does not list them one by one here.
[0245] Finally, it should be noted that the various embodiments in this specification are described in a progressive or parallel manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices, terminal devices, and systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description of the method part.
[0246] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0247] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the core idea or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A traffic sign detection method, characterized in that, The method includes: Obtaining a traffic real - scene image in the driving direction of the vehicle; Performing compression processing on the traffic real - scene image to obtain an image to be detected; the resolution of the image to be detected is less than the resolution of the traffic real - scene image; Performing main traffic sign detection on the image to be detected, and based on the main traffic sign detection result, obtaining a potential image area where there is a first auxiliary traffic sign that matches the detected first main traffic sign; the potential image area includes a local image area of the traffic real - scene image; Inputting the potential image area into an auxiliary traffic sign detection network to obtain the first auxiliary traffic sign included in the potential image area and the area position information of the first auxiliary traffic sign in the potential image area; Obtaining the potential area position information of the potential image area in the traffic real - scene image; Determining the image filling data of the traffic real - scene image and the scaling ratio of the local image area where the first auxiliary traffic sign exists obtained from the traffic real - scene image during the process of obtaining the potential image area; wherein, the image filling data includes the filling distances in different diffusion directions; the potential image area is obtained by compressing or magnifying the local image area; Using the filling distances in different diffusion directions, the scaling ratio, and the potential area position data to perform reduction processing on the obtained area position information to obtain the target position information of the first auxiliary traffic sign in the traffic real - scene image.
2. The method according to claim 1, wherein The performing main traffic sign detection on the image to be detected, and based on the main traffic sign detection result, obtaining a potential image area where there is a first auxiliary traffic sign that matches the detected first main traffic sign includes: Inputting the image to be detected into a main traffic sign detection network to obtain the first main traffic sign included in the traffic real - scene image; Taking the first main traffic sign included in the traffic real - scene image as a reference object to determine a local image area where there is a first auxiliary traffic sign that matches the first main traffic sign; Forming a potential image area from the determined local image areas.
3. The method according to claim 2, wherein The taking the first main traffic sign included in the traffic real - scene image as a reference object to determine a local image area where there is a first auxiliary traffic sign that matches the first main traffic sign includes: In the traffic real - scene image, taking the obtained first main traffic sign as the center and spreading a preset distance around to obtain a local image area where there is a first auxiliary traffic sign that matches the first main traffic sign.
4. The method according to claim 3, characterized in that, During the process of spreading a preset distance around with the obtained first main traffic sign as the center, the method further includes: Detecting that the first boundary distance between the first main traffic sign and the first image boundary is less than the preset distance; the first image boundary refers to the boundary of the traffic real - scene image in the first diffusion direction; Spreading from the first main traffic sign to the first image boundary in the first diffusion direction and continuing to perform image filling and spreading according to the first image filling method until the diffusion distance in the first diffusion direction reaches the preset distance; The potential image region formed by the determined local image regions includes: The local image region and the image filling region in the obtained traffic real-scene image form a potential image region where the first auxiliary traffic sign exists.
5. The method according to claim 2, wherein Determining the local image region where the first auxiliary traffic sign that matches the first main traffic sign exists, with reference to the first main traffic sign included in the traffic real-scene image, includes: According to the sign type of the first main traffic sign included in the traffic real-scene image, obtaining the first position deployment relationship corresponding to the main traffic sign of the sign type; wherein, the first position deployment relationship refers to the relative position relationship between the main traffic sign of the sign type and the auxiliary traffic sign that matches the main traffic sign in traffic road facilities. According to the first position deployment relationship, determine the local image region where the first auxiliary traffic sign that matches the first main traffic sign exists from the traffic real-scene image.
6. The method according to claim 2, wherein The forming of the potential image region by the determined local image regions includes: Comparing the area of the determined local image region with a preset auxiliary traffic sign detection area. If the area of the local image region is smaller than the preset auxiliary traffic sign detection area, perform magnification processing on the local image region to obtain a potential image region with the preset auxiliary traffic sign detection area. If the area of the local image region is larger than the preset auxiliary traffic sign detection area, perform compression processing on the local image region to obtain a potential image region with the preset auxiliary traffic sign detection area. Wherein, the resolution of the potential image region is smaller than the resolution of the traffic real-scene image.
7. The method according to claim 1, characterized in that, If the number of the first main traffic signs is multiple, obtaining the potential image region where the first auxiliary traffic sign that matches the detected first main traffic sign exists includes: According to the detected multiple first main traffic signs, obtain the potential image regions where the first auxiliary traffic signs that respectively match the multiple first main traffic signs exist; wherein, the potential image region includes the corresponding first auxiliary traffic sign and at least part of the first main traffic sign that matches the first auxiliary traffic sign. The detecting of the auxiliary traffic sign for the potential image region includes: Detecting the auxiliary traffic sign for each obtained potential image region.
8. The method according to any one of claims 1 to 6, characterized in that The method further includes: Rendering the traffic real-scene image according to the target position information. Outputting the rendered navigation image, and displaying the detected first main traffic sign and the first auxiliary traffic sign that matches the first main traffic sign in the navigation image; or, Popping up a traffic prompt window in the navigation image, and displaying traffic prompt information for the first auxiliary traffic sign in the traffic prompt window; the traffic prompt information is determined based on the auxiliary traffic indication information of the first auxiliary traffic sign, and the driving state information and / or system time of the vehicle; or, Performing voice broadcast on the traffic prompt information.
9. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the auxiliary traffic indication information of the corresponding first auxiliary traffic sign according to the target position information; Utilize the target position information and the auxiliary traffic indication information to obtain the structured perception result of the corresponding first auxiliary traffic sign in the traffic real-scene image; Report the structured perception result to the map service device, and the map service device updates the corresponding map data by using the structured perception result.
10. The method according to any one of claims 1 to 6, characterized in that, The method further includes: According to the target position information, obtain the relative distance between the vehicle and the corresponding first auxiliary traffic sign; Obtain the main traffic indication information of the first main traffic sign included in the traffic real-scene image, and the auxiliary traffic indication information of the first auxiliary traffic sign matched with the first main traffic sign; Generate a road guidance message in the driving direction of the vehicle by using the relative distance, the main traffic indication information, and the auxiliary traffic indication information; Perform voice broadcast on the road guidance message.
11. A traffic sign detection device, characterized in that, The device includes: A traffic real-scene image acquisition module, configured to acquire a traffic real-scene image in the driving direction of the vehicle; A to-be-detected image obtaining module, configured to perform compression processing on the traffic real-scene image to obtain a to-be-detected image; the resolution of the to-be-detected image is less than the resolution of the traffic real-scene image; A potential image area obtaining module, configured to perform main traffic sign detection on the to-be-detected image, and obtain a potential image area where a first auxiliary traffic sign that matches the detected first main traffic sign exists according to the main traffic sign detection result; the potential image area includes a partial image area of the traffic real-scene image; A target position information acquisition module, configured to perform auxiliary traffic sign detection on the potential image area, and obtain the target position information of the detected first auxiliary traffic sign mapped in the traffic real-scene image according to the auxiliary traffic sign detection result; The target position information acquisition module includes: An auxiliary traffic sign detection unit, configured to input the potential image area into an auxiliary traffic sign detection network to obtain the first auxiliary traffic sign included in the potential image area, and the regional position information of the first auxiliary traffic sign in the potential image area; A potential regional position information acquisition unit, configured to obtain the potential regional position information of the potential image area in the traffic real-scene image; A coordinate conversion processing unit, configured to perform coordinate conversion processing on the regional position information according to the obtaining manner of the potential image area and the potential regional position information to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image; The coordinate conversion processing unit includes: A first data determination unit, configured to determine the image padding data of the traffic real-scene image and the scaling ratio of the partial image area where the first auxiliary traffic sign exists obtained from the traffic real-scene image during the process of obtaining the potential image area; Wherein, the image padding data includes the padding distances in different diffusion directions; the potential image area is obtained by compressing or magnifying the partial image area; The first position restoration processing unit is further configured to use the filling distances in different diffusion directions, the scaling ratio, and the potential area position data to perform restoration processing on the obtained area position information, so as to obtain the target position information of the first auxiliary traffic sign in the traffic real-scene image.
12. The device according to claim 11, characterized in that, The potential image area obtaining module includes: The main traffic sign detection unit is configured to input the image to be detected into the main traffic sign detection network to obtain the first main traffic sign included in the traffic real-scene image; The local image area determination unit is configured to use the first main traffic sign included in the traffic real-scene image as a reference to determine the local image area where the first auxiliary traffic sign matching the first main traffic sign exists; The potential image area composition unit is configured to compose the potential image area from the determined local image area.
13. The device according to claim 12, characterized in that, The local image area determination unit includes: The first diffusion processing unit is configured to diffuse a preset distance around the obtained first main traffic sign in the traffic real-scene image to obtain the local image area where the first auxiliary traffic sign matching the first main traffic sign exists.
14. The device according to claim 13, characterized in that The local image area determination unit includes: The first detection unit is configured to detect that the first boundary distance between the first main traffic sign and the first image boundary is less than the preset distance during the process of diffusing a preset distance around the obtained first main traffic sign; the first image boundary refers to the boundary of the traffic real-scene image in the first diffusion direction; The first filling unit is configured to diffuse from the first main traffic sign to the first image boundary in the first diffusion direction and continue to perform image filling diffusion according to the first image filling method until the diffusion distance in the first diffusion direction reaches the preset distance; The potential image area composition unit is specifically configured to: compose the potential image area where the first auxiliary traffic sign exists from the obtained local image area and the image filling area in the traffic real-scene image.
15. The device according to claim 12, characterized in that, The local image area determination unit includes: The first position deployment relationship obtaining unit is configured to obtain the first position deployment relationship corresponding to the main traffic sign of the sign type according to the sign type of the first main traffic sign included in the traffic real-scene image; wherein, the first position deployment relationship refers to the relative position relationship between the main traffic sign of the sign type and the auxiliary traffic sign matching the main traffic sign in the traffic road facilities; The second diffusion processing unit is configured to determine the local image area where the first auxiliary traffic sign matching the first main traffic sign exists from the traffic real-scene image according to the first position deployment relationship.
16. The device according to claim 12, characterized in that, The potential image area composition unit includes: The comparison unit is configured to compare the area of the determined local image area with the preset auxiliary traffic sign detection area; The first obtaining unit is configured to, when the area of the local image area is less than the preset auxiliary traffic sign detection area, perform magnification processing on the local image area to obtain the potential image area with the preset auxiliary traffic sign detection area; A second obtaining unit, configured to perform compression processing on the local image region when the area of the local image region is greater than the preset auxiliary traffic sign detection area, so as to obtain a potential image region having the preset auxiliary traffic sign detection area; Wherein, the resolution of the potential image region is less than the resolution of the traffic real-scene image.
17. The device according to claim 11, characterized in that, When the number of detected first main traffic signs is multiple, the potential image region obtaining module is specifically configured to obtain potential image regions where the respective first auxiliary traffic signs matching the multiple detected first main traffic signs exist, based on the multiple detected first main traffic signs; wherein, the potential image region includes the corresponding first auxiliary traffic sign and at least part of the first main traffic sign matching the first auxiliary traffic sign; Correspondingly, the above-mentioned auxiliary traffic sign detection unit performs auxiliary traffic sign detection on each of the obtained potential image regions.
18. The device according to any one of claims 11-16, characterized in that, The device further includes: An image rendering module, configured to render the traffic real-scene image according to the target position information; An image output module, configured to output the rendered navigation image, and display the detected first main traffic signs and the first auxiliary traffic signs matching the first main traffic signs in the navigation image; or, Pop up a traffic prompt window in the navigation image, and display traffic prompt information for the first auxiliary traffic sign in the traffic prompt window; the traffic prompt information is determined based on the auxiliary traffic indication information of the first auxiliary traffic sign, and the driving state information and / or system time of the vehicle; or, Perform voice broadcast on the traffic prompt information.
19. The device according to any one of claims 11-16, characterized in that, The device further includes: An auxiliary traffic indication information obtaining module, configured to obtain the auxiliary traffic indication information of the corresponding first auxiliary traffic sign according to the target position information; A structured perception result obtaining module, configured to obtain a structured perception result of the corresponding first auxiliary traffic sign in the traffic real-scene image by using the target position information and the auxiliary traffic indication information; A structured perception result reporting module, configured to report the structured perception result to a map service device, and the map service device updates the corresponding map data by using the structured perception result.
20. The device according to any one of claims 11-16, characterized in that, The above-mentioned device includes: A distance obtaining module, configured to obtain the relative distance between the vehicle and the corresponding first auxiliary traffic sign according to the target position information; A traffic indication information acquisition module, configured to acquire the main traffic indication information of the first main traffic sign included in the traffic real-scene image and the auxiliary traffic indication information of the first auxiliary traffic sign matching the first main traffic sign; A road guidance message generation module, configured to generate a road guidance message in the driving direction of the vehicle by using the relative distance, the main traffic indication information, and the auxiliary traffic indication information; A road guidance message broadcast module, configured to perform voice broadcast on the road guidance message.
21. A terminal device, characterized in that, The terminal device includes: A communication interface; A memory for storing a program for implementing the traffic sign detection method according to any one of claims 1-10; A processor for loading and executing the program stored in the memory to implement the steps of the traffic sign detection method according to any one of claims 1-10.
22. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is called and executed by a processor to implement the traffic sign detection method according to any one of claims 1-10.
23. A computer program product, characterized in that, The computer program product includes instructions that, when run on a computer device, cause the computer device to execute the traffic sign detection method according to any one of claims 1-10.
Citation Information
Patent Citations
A traffic sign detection method and system based on step-by-step deep learning
CN109815906A
Traffic sign board information acquisition method and system for high-precision map production
CN110501018A
Road sign recognition device
CN112740293A
KR20200069910A