An active traffic sign recognition method and device based on fusion of multiple cameras
By using a fusion recognition method combining panoramic and telephoto cameras, the problem of simultaneously recognizing multiple traffic signs at a distance was solved, achieving efficient traffic sign detection and recognition and improving recognition performance.
Patent Information
- Application Number
- CN202510322515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In existing technologies, telephoto lenses cannot simultaneously image traffic signs in multiple directions within a scene. Schemes based on small target detection result in poor features and excessive noise, making it impossible to effectively identify multiple traffic signs.
An active traffic sign recognition method that integrates panoramic and telephoto cameras is adopted. The panoramic camera initially identifies multiple signs, tracks signs with low confidence, adjusts the telephoto camera angle to acquire high-definition images, and performs data fusion recognition.
It achieves efficient recognition of multiple traffic signs, expands the detection and recognition range, improves recognition performance, and combines the wide field of view of a panoramic camera with the high-definition imaging advantages of a telephoto camera.
Smart Images

Figure CN120375323B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the field of intelligent vehicles, and in particular, to an active traffic sign recognition method and device based on fusion of multiple cameras. BACKGROUND
[0002] Currently, traffic sign recognition is one of the important tasks of intelligent vehicle visual perception. Traffic signs are traffic facilities that convey warning, prohibition, indication and road guidance information through symbols and text. As one of the traffic participants, intelligent vehicles must be able to accurately recognize these traffic signs and follow the instructions of the signs to constrain their own behavior. Thus, intelligent vehicles can only be integrated into increasingly complex traffic scenarios. When the intelligent vehicle is traveling at a certain speed, the earlier the meaning of the traffic sign is recognized, the more reaction time the intelligent vehicle can have to make judgments about the current traffic situation and make decisions and plans. Therefore, how to accurately detect and recognize traffic signs at a long distance has important research value in the field of intelligent vehicles.
[0003] For the detection and recognition of traffic signs at a long distance, there are two different solutions: one is a small target detection-based solution. This solution does not change the input image, but improves the network structure to improve the recognition performance of the algorithm for small targets. This solution mostly focuses on improving the neural network to extract multi-scale invariant features of the target, and completes the detection and classification of traffic signs based on these features. However, when the number of target pixels is too small, the features of the target will be poor and the noise will be high, so the small target detection-based solution has very limited improvement on the detection result. The other is to use a telephoto lens to improve the imaging quality of the target at a long distance to improve the detection and recognition effect of the target. Since the field of view of the telephoto lens is usually small, it cannot simultaneously image multiple traffic signs in different directions in the scene.
[0004] Therefore, how to simultaneously image and recognize multiple traffic signs in different directions in the scene is a problem that needs to be solved at present. SUMMARY
[0005] To solve the problem that the imaging quality of the target at a long distance cannot be improved by using a telephoto lens to simultaneously image multiple traffic signs in different directions in the scene, and the small target-based traffic sign detection results in poor features and high noise, the present specification provides an active traffic sign recognition method and device based on fusion of multiple cameras. Through the active traffic sign recognition algorithm based on fusion of a panoramic camera and a long-focus camera, the panoramic camera and the long-focus camera are efficiently combined, the recognition performance is greatly improved, and the range of traffic sign detection and recognition is effectively expanded.
[0006] To solve any of the above technical problems, the specific technical solutions of the embodiments of the present specification are as follows:
[0007] In one aspect, the embodiments of the present specification provide an active traffic sign recognition method based on fusion of multiple cameras, comprising:
[0008] In vehicle driving, a panoramic camera collects panoramic images and performs preliminary identification, identifying a plurality of signs in the panoramic images;
[0009] Tracking the plurality of signs whose relative positions change due to the driving of the vehicle, judging the confidence of the plurality of signs identified preliminarily, and inputting the signs with a confidence lower than a preset threshold to a sorting module;
[0010] The sorting module calculates an importance indicator of each sign with a confidence lower than a preset threshold, sorts according to the importance indicator, and determines the sign with the highest importance indicator as the current target sign;
[0011] According to the changing relative position of the current target sign in the driving of the vehicle, adjusting the target angle of the long-focus camera, and collecting the high-definition image of the current target sign located at the target angle, the next sign is sorted as the current target sign according to the importance indicator, and the high-definition image of the current target sign is collected in turn;
[0012] High-definition identification is performed on the high-definition image, the results of the high-definition identification are data-fused with the results of the preliminary identification, and a traffic sign recognition result is generated.
[0013] Further, judging the confidence of the plurality of signs identified preliminarily further comprises:
[0014] The sign with a confidence higher than a preset threshold directly generates the traffic sign recognition result.
[0015] Further, the sorting module calculating an importance indicator of each sign with a confidence lower than a preset threshold further comprises:
[0016] Identifying the type of the sign, the type at least including: warning sign, prohibition sign, indication sign or guide sign;
[0017] Obtaining the current driving task of the vehicle, and calculating the importance indicator of the sign according to a preset rule and the type of the traffic sign.
[0018] Further, according to the changing relative position of the current target sign in the driving of the vehicle, adjusting the target angle of the long-focus camera further comprises:
[0019] judging an initial coordinate of a center point of the current target sign in the panoramic image;
[0020] judging an angle deviation according to a driving state of the vehicle, and generating the target angle according to the initial coordinate and the angle deviation;
[0021] adjusting the long-focus camera according to the target angle, and adjusting a focal length of the long-focus camera according to a focus evaluation of the panoramic image.
[0022] Further, adjusting the focal length of the long-focus camera according to the focus evaluation of the panoramic image further comprises:
[0023] obtaining a gray value of the panoramic image, and calculating a focus evaluation function according to the gray value;
[0024] calculating a target focal length when the focus evaluation function is greater than a preset threshold;
[0025] taking the target focal length as the focal length of the long-focus camera.
[0026] Further, performing high-definition recognition on the high-definition image, and performing data fusion on the high-definition recognition result and the preliminary recognition result to generate a traffic sign recognition result further comprises:
[0027] judging a number of traffic signs in the high-definition recognition result;
[0028] if the high-definition recognition result contains only one traffic sign, taking a type and a confidence degree of the traffic sign as the traffic sign recognition result;
[0029] if the high-definition recognition result contains multiple traffic signs, performing data fusion on the high-definition recognition result and the preliminary recognition result to generate a traffic sign recognition result.
[0030] Further, performing data fusion on the high-definition recognition result and the preliminary recognition result to generate a traffic sign recognition result further comprises:
[0031] saving the preliminary recognition result in the panoramic image, and calculating center coordinates of multiple traffic signs in the panoramic image;
[0032] generating a virtual viewfinder according to the center coordinates and the panoramic image;
[0033] calculating feature vectors of multiple traffic signs in the high-definition image and the virtual viewfinder respectively, and matching traffic signs in the panoramic image with traffic signs in the high-definition image;
[0034] The matching result, the type corresponding to each traffic sign, and the confidence are taken as the recognition result of the traffic sign.
[0035] In another aspect, the embodiments of the present specification also provide an active traffic sign recognition device based on fusion of multiple cameras, which comprises:
[0036] A preliminary recognition unit is configured to collect panoramic images by a vehicle-mounted panoramic camera and perform preliminary recognition during vehicle driving, and identify multiple signs in the panoramic images.
[0037] A confidence tracking unit is configured to track the multiple signs whose relative positions are constantly changing due to vehicle driving, judge the confidence of the multiple signs in the preliminary detection, and input the signs with a confidence lower than a preset threshold to a sorting module.
[0038] An importance sorting unit comprises the sorting module, configured to calculate an importance indicator of each sign with a confidence lower than a preset threshold, sort according to the importance indicator, and determine the sign with the highest importance indicator as a current target sign.
[0039] A high-definition collection unit is configured to adjust a target angle of a long-focus camera according to the changing relative position of the current target sign during vehicle driving, collect a high-definition image of the current target sign at the target angle, sort the next sign as the current target sign according to the importance indicator, and sequentially collect high-definition images of the current target signs.
[0040] A high-definition recognition unit is configured to perform high-definition recognition on the high-definition images, perform data fusion on the recognition result of the high-definition recognition and the result of the preliminary detection, and generate a traffic sign recognition result.
[0041] In another aspect, the embodiments of the present specification also provide a computer device, which comprises a memory, a processor, and a computer program stored in the memory, and the processor implements the above method when executing the computer program.
[0042] In another aspect, the embodiments of the present specification also provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above method.
[0043] With the embodiments of the present specification, during the driving of the vehicle, the vehicle-mounted panoramic camera collects panoramic images and performs preliminary identification, identifies a plurality of signs in the panoramic images, and tracks the signs during the driving of the vehicle, so as to judge the confidence of each sign, and input the signs with low confidence into a sorting module, so as to further confirm whether the identified result is correct; in the sorting module, the importance index of each sign is calculated, and the signs are sorted according to the importance index, and then the high-definition images of each sign are collected in sequence according to the sorting, wherein during the collection, the target angle of the long-focus camera needs to be adjusted according to the relative position change of the target sign during the driving of the vehicle, after the high-definition images of each sign are obtained, the high-definition images are subjected to high-definition identification, the result of the high-definition identification is data-fused with the result of the preliminary detection, and a traffic sign recognition result is generated. The efficient combination of the panoramic camera and the long-focus camera during the traffic sign recognition is realized, the recognition performance is greatly improved, and the range of traffic sign detection and recognition is effectively expanded. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can also be obtained by those skilled in the art without creative labor.
[0045] Figure 1 An implementation system schematic diagram of a kind of active traffic sign recognition method based on multiple camera fusion in the present specification embodiment is shown;
[0046] Figure 2 A flowchart of a kind of active traffic sign recognition method based on multiple camera fusion in the present specification embodiment is shown;
[0047] Figure 3 A flowchart of adjusting the target angle of long-focus camera in the present specification embodiment is shown;
[0048] Figure 4 A flowchart of adjusting the focal length of the long-focus camera according to the focusing evaluation of the panoramic image in the present specification embodiment is shown;
[0049] Figure 5 A flowchart of generating traffic sign recognition result in the present specification embodiment is shown;
[0050] Figure 6 A flowchart of generating traffic sign recognition result for multiple target signs in the present specification embodiment is shown;
[0051] Figure 7Fig. 1 shows a structural schematic diagram of an active traffic sign recognition device based on camera fusion according to an embodiment of the present specification.
[0052] Figure 8 Fig. 2 shows a structural schematic diagram of a computer device according to an embodiment of the present specification.
[0053]
Explanation of reference signs
[0054] 101, processor
[0055] 102, panoramic camera
[0056] 103, long-focus camera
[0057] 701, preliminary recognition unit
[0058] 702, confidence tracking unit
[0059] 703, importance ranking unit
[0060] 704, high-definition acquisition unit
[0061] 705, high-definition recognition unit
[0062] 802, computer device
[0063] 804, processing device
[0064] 806, storage resource
[0065] 808, driving mechanism
[0066] 810, input / output module
[0067] 812, input device
[0068] 814, output device
[0069] 816, presentation device
[0070] 818, graphical user interface
[0071] 820, network interface
[0072] 822, communication link
[0073] 824, communication bus DETAILED DESCRIPTION
[0074] With reference to the accompanying drawings, the technical solutions in the embodiments of the present specification will be described clearly and completely. Obviously, the described embodiments are only part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the embodiments of the present specification, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present specification.
[0075] It should be noted that the terms "first", "second" and the like in the description and claims of the present specification and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present specification described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or apparatus including a series of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or apparatus.
[0076] It should be noted that the acquisition, storage, use, processing and the like of data in the technical solutions of the present specification comply with the relevant provisions of national laws and regulations.
[0077] It should be noted that in the present specification, some software, components, models and the like may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the technical solutions in the present application, but it does not mean that the applicant has or will necessarily use the scheme.
[0078] As Figure 1 shown is an implementation system schematic diagram of an active traffic sign recognition method based on multiple camera fusion in the embodiments of the present specification, including a processor 101, a panoramic camera 102 and a long-focus camera 103. The processor 101 and the panoramic camera 102 and the long-focus camera 103 can communicate through a network, which can include a bus network, a local area network (Local Area Network, LAN for short), a wide area network (Wide Area Network, WAN for short), the Internet or a combination thereof, and is connected to a website, a user device (such as a computing device) and a backend system.
[0079] Wherein, the panoramic camera 102 can transmit the panoramic image to the processor 101, the processor 101 identifies the mark in the panoramic image, then analyzes it, identifies the mark with insufficient confidence, enables the long-focus camera 103 to collect the high-definition image of the same mark again, and fuses the panoramic image and the high-definition image to generate the traffic sign recognition result. In addition, it needs to be explained that, Figure 1 The application environment shown is only one application environment provided by the embodiment of the present specification, and other scenes can also be included in actual application, which is not limited by the present specification.
[0080] In view of the problems in the prior art, the embodiment of the present specification provides an active traffic sign recognition algorithm based on the fusion of panoramic cameras and long-focus cameras, which realizes the efficient combination of panoramic cameras and long-focus cameras, greatly improves the recognition performance, and further effectively expands the range of traffic sign detection and recognition.
[0081] Figure 2 The flowchart of the active traffic sign recognition method based on the fusion of multiple cameras is shown. The process of recognizing traffic signs by multiple camera fusion is described in the figure. The order of the steps listed in the embodiment is only one of the many step execution orders, and does not represent the only execution order. When the system or device product is executed in practice, it can be executed in sequence or in parallel according to the method shown in the embodiment or the drawing.
[0082] Specifically, as Figure 2 As shown, the method can be executed by the processor 101, and the method can include:
[0083] Step 201: In the vehicle driving, the vehicle-mounted panoramic camera collects the panoramic image and performs preliminary identification, and identifies a plurality of marks in the panoramic image;
[0084] Step 202: Track the plurality of marks whose relative positions are constantly changing due to the vehicle driving, judge the confidence of the plurality of marks identified preliminarily, and input the mark with a confidence lower than a preset threshold to a sorting module;
[0085] Step 203: The sorting module calculates the importance index of each mark with a confidence lower than a preset threshold, sorts according to the importance index, and determines the mark with the highest importance index as the current target mark;
[0086] Step 204: According to the changing relative position of the current target mark in the vehicle driving, adjust the target angle of the long-focus camera, collect the high-definition image of the current target mark located at the target angle, and sort the next mark as the current target mark according to the importance index, and sequentially collect the high-definition image of the current target mark;
[0087] Step 205: high-definition recognition is performed on the high-definition image, and the result of the high-definition recognition is data-fused with the result of the preliminary recognition to generate a traffic sign recognition result.
[0088] With the embodiments of the present specification, in the process of vehicle driving, a vehicle-mounted panoramic camera collects panoramic images and performs preliminary recognition, recognizes a plurality of signs in the panoramic images, and tracks these signs in the process of vehicle driving, so as to judge the confidence of each sign, input the sign with low confidence into a sorting module, so as to further confirm whether the recognition result is correct; in the sorting module, the importance index of each sign is calculated, and the signs are sorted according to the importance index, and then the high-definition image of each sign is collected in sequence according to the sorting, wherein in the process of collection, the target angle of the long-focus camera needs to be adjusted according to the relative position change of the target sign in the process of vehicle driving, after the high-definition image of each sign is obtained, high-definition recognition is performed on the high-definition image, and the result of the high-definition recognition is data-fused with the result of the preliminary detection to generate a traffic sign recognition result. The efficient combination of the panoramic camera and the long-focus camera in the process of traffic sign recognition is realized, the recognition performance is greatly improved, and the range of traffic sign detection and recognition is effectively expanded.
[0089] In an embodiment of the present specification, a panoramic camera collects panoramic images of the surrounding scene, and then inputs the panoramic images to a traffic sign detection and recognition module based on the panoramic images to perform preliminary detection and recognition of traffic signs, and then inputs the detection and recognition result to a Deep SORT module for tracking; the Deep SORT module is used to track all signs and judge the confidence thereof, if the confidence of a sign is higher than a threshold value, it is indicated that the tracked object is well detected and does not need to be further processed, and the panoramic image can be directly detected and recognized, if the confidence of part of the signs is lower than the threshold value, it is indicated that these signs need to be high-definition recognized by a long-focus camera. The sign needing to be high-definition recognized by the long-focus camera is input to a sorting module, and the importance index of each detection target is calculated and the order of high-definition recognition is determined according to the importance index, the detection target with the highest importance index is set as the current target sign, and then the angle of the target sign relative to the panoramic camera is calculated as the initial target angle of the long-focus camera. The long-focus camera will first rotate to the initial target angle, and then control the angle and focal length to collect the high-definition image of the target, and the high-definition image is input to a target detection and recognition module based on the high-definition image to perform high-precision detection and recognition.
[0090] In another embodiment of the present specification, judging the confidence of the plurality of signs preliminarily recognized further comprises:
[0091] The sign with the confidence higher than the preset threshold value directly generates the traffic sign recognition result.
[0092] In the embodiment of the present specification, in order to obtain the importance index of each sign to sort the collection sequence of the high-definition image, the sorting module further comprises:
[0093] identifying the type of the sign, the type at least including: warning sign, prohibition sign, indication sign or guide sign;
[0094] obtaining the current driving task of the vehicle, and calculating the importance index of the sign according to the preset rule and the type of the traffic sign.
[0095] Specifically, the multi-target sorting method based on importance comprehensively considers the confidence of the panoramic detection result, the running state of the vehicle, the relative angle between the traffic sign and the vehicle, and the importance of the traffic sign itself, and constructs an index for evaluating the importance of each target to be detected in the current state based on the above factors. Its beneficial effects are that based on this index, multiple targets to be detected can be sorted, and finally the uncertain targets are sequentially arranged according to the size of importance.
[0096] Exemplarily, based on the setting principle of the traffic sign, in Deep SORT, more attention should be paid to the warning sign and the prohibition sign, and the importance of each type of traffic sign is modeled by using a sigmoid function. It is assumed that the importance of the warning sign is S(3)=0.95, the importance of the prohibition sign is S(2)=0.88, the importance of the indication sign is S(1)=0.73, and the importance of the guide sign is S(0)=0.5. At the same time, the importance of the sign type is also related to the confidence of the recognition result. The higher the confidence, the more acceptable the evaluation based on the sign type is. Therefore, the importance based on the sign type is defined as: κ1=α0S(x), wherein κ1 is the importance based on the sign type, and α0 is the confidence of the target sign to be detected.
[0097] In another embodiment of the present specification, the importance of the traffic sign is not only related to the type of the sign itself, but also related to the driving state of the vehicle. When the vehicle is driving straight, the sign in front is more important, and when the vehicle is turning, the sign in the turning direction is more important. Therefore, the angle between the preview point and the target sign can be used as an importance evaluation index.
[0098] In the embodiment of the present specification, in order to obtain the high-definition image of the target sign using the long-focus camera, as shown in Figure 3 adjusting the target angle of the long-focus camera according to the relative position of the current target sign changing in the driving of the vehicle further comprises:
[0099] Step 301: judging the initial coordinates of the center point of the current target sign in the panoramic image;
[0100] Step 302: judging the angle deviation according to the driving state of the vehicle, and generating the target angle according to the initial coordinate and the angle deviation;
[0101] Step 303: adjusting the long-focus camera according to the target angle, and adjusting the focal length of the long-focus camera according to the focus evaluation of the panoramic image.
[0102] Specifically, first, the initial coordinate of the center point of the current target sign in the panoramic image is judged. However, during the driving of the vehicle, the traffic sign is fixed in position, but due to the movement of the vehicle, the camera is also moved, and relative movement is formed between the traffic sign and the camera, which causes the traffic sign to deviate from the center of the image and produce deviation in the horizontal and vertical directions. Therefore, feedback control is performed based on the pixel deviation of the actual position of the target sign from the center of the image to control the long-focus camera to adjust to the target angle.
[0103] In the embodiments of the present application, in order to use the long-focus camera to obtain a high-definition image of the target sign, after the long-focus camera is controlled to adjust to the target angle, the focal length of the long-focus camera controlled by the depth of focus method can be used to transform the focal length, collect a series of target images with different degrees of blurring, calculate the focus evaluation function of each focal length, and find the focal length with the highest image clarity, i.e. the best imaging focal length. As shown in Figure 4 Further adjusting the focal length of the long-focus camera according to the focus evaluation of the panoramic image comprises:
[0104] Step 401: obtaining the gray value of the panoramic image, and calculating the focus evaluation function according to the gray value;
[0105] Step 402: when the focus evaluation function is greater than a preset threshold, calculating the target focal length;
[0106] Step 403: taking the target focal length as the focal length of the long-focus camera.
[0107] Exemplarily, the focal length control of the long-focus camera uses the depth of focus method. The core of the depth of focus method is to design a suitable focus evaluation function. The information entropy of the image can be used as the focus evaluation function, thereby effectively indicating the clarity of the image. When the image tends to be out of focus, the gray value tends to be single, the information content is small, and the information entropy is small. When the image tends to be in focus, the diversity of the gray of the image becomes large, and the information entropy also becomes large.
[0108] wherein the gray entropy E of the image can be represented as: A
[0109]
[0110] wherein P i,j P is a spatial feature of gray scale, denoted as P i,j = f(i, j) / (W A H A ); wherein (i, j) is a feature pair, i is a gray scale value of the image, j is a neighborhood gray scale mean value, f(i, j) is a probability of the feature pair (i, j) appearing, W A is an image width of the long-focus camera, and H A is an image height of the long-focus camera. Adjustment of the focal length of the long-focus camera is realized by the depth-of-focus method.
[0111] In an embodiment of the present specification, after successfully collecting a high-definition image of the target sign, in order to generate a final traffic sign recognition result, the high-definition image is subjected to high-definition recognition, the result of the high-definition recognition is data-fused with the result of the preliminary recognition, and a traffic sign recognition result is generated, which further comprises: Figure 5
[0112] Step 501: judging the number of traffic signs in the result of the high-definition recognition;
[0113] Step 502: if the result of the high-definition recognition contains only one traffic sign, the type and confidence thereof are taken as the recognition result of the traffic sign;
[0114] Step 503: if the result of the high-definition recognition contains multiple traffic signs, the result of the high-definition recognition is data-fused with the preliminary recognition result to generate a traffic sign recognition result.
[0115] Specifically, the feature vectors of the targets on the high-definition image and the virtual image need to be calculated respectively and input to a matching module for matching, so that the preliminary recognition result on the panoramic image is corresponded to the result of the high-definition recognition. In this way, the accurate detection and recognition result on the high-definition image can be assigned to the target on the panoramic image for tracking and management in the whole life cycle. The whole life cycle refers to the whole process from the first detection of the target to the disappearance of the target on the panoramic image. In an embodiment of the present specification, a virtual camera model equivalent to the long-focus camera can be used to find the corresponding part of the long-focus camera from the detection result of the panoramic image for matching, thereby reducing the matching with irrelevant targets and improving the accuracy and efficiency of the matching.
[0116] In an embodiment of the present specification, in some cases, multiple traffic signs in the result of the high-definition recognition will appear multiple targets to be detected, and at this time, the matching between these targets is needed to accurately data-fuse the detection result of the long-focus camera with the detection result of the panoramic image. In order to obtain multiple sign recognition results in the high-definition recognition, as shown in Figure 6 The result of the high-definition identification is fused with the preliminary identification result to generate a traffic sign identification result, and the traffic sign identification result further includes:
[0117] Step 601: saving the preliminary identification result in the panoramic image and calculating the center coordinates of the multiple traffic signs in the panoramic image;
[0118] Step 602: generating a virtual viewfinder according to the center coordinates and the panoramic image;
[0119] Step 603: calculating the feature vectors of the multiple traffic signs in the high-definition image and the virtual viewfinder respectively, and matching the traffic signs in the panoramic image with the traffic signs in the high-definition image;
[0120] Step 604: taking the matching result, the type corresponding to each traffic sign, and the confidence as the identification result of the traffic sign.
[0121] In the embodiments of the present specification, after the traffic sign in the panoramic image is fused with the corresponding high-definition image, the identification result of the traffic sign can be output, the traffic sign can be tracked according to the identification result, and the vehicle can be controlled to change the driving state according to the traffic sign. Because the generation order of the result of high-definition identification is determined according to the importance index, the identification result generated first is definitely the sign type with the highest importance index, for example, the identification result of a warning sign, so that whether to change the driving state of the vehicle is determined according to the warning sign first, and the efficiency of identification is improved.
[0122] Specifically, the angle and focal length of the long-focus camera corresponding to the high-definition image are fed back to the data fusion module; the panoramic camera is input to generate a virtual viewfinder corresponding to the high-definition image. Because the virtual viewfinder is derived from the panoramic image, it will inherit the detection result of the panoramic image; the detection result on the virtual viewfinder is matched with the detection result on the high-definition image. Thus, the correspondence between the multiple detection results on the high-definition image and the targets on the panoramic image is calculated; the sign type and confidence of each high-definition identification result are respectively assigned to the corresponding targets on the panoramic image, so as to realize the tracking and management of the whole life cycle of the multiple target signs in the panoramic image, thereby improving the accuracy and detection efficiency of data fusion.
[0123] Exemplarily, the virtual viewfinder is denoted as A picture, and the high-definition image corresponding to A and collected by the long-focus camera is denoted as B picture. It is assumed that there are n frame-selected targets in the A picture and m frame-selected targets in the B picture, which are denoted as a i ,i=1,...,n and b j ,j=1,...,m respectively.
[0124] Extract each bounding target a using the feature extraction network in Deep SORT. i and b j eigenvector f a (i) and f b (i);
[0125] Calculate each target a in image A i With each target b in image B j To determine the similarity, we construct a matching matrix M. Assuming the similarity score function is s(·,·), the elements of the matching matrix can be represented as:
[0126] M i,j =s(f a (i),f b (i)),
[0127] Among them, M i,j Let represent the element in the i-th row and j-th column of the matching matrix M, and let s(·,·) be the similarity score function.
[0128] Defined as:
[0129] In the formula, x and y are two vectors, and n V Let be the dimension of the vector. Chi-square distribution detection is used to measure the reliability of the match; for each target a in image A... i Calculate its similarity score with the target in all B images, and obtain a score vector v. i Then v i =[s(f a (i),f b (1)),s(f a (i),f b (2)),...,s(f a (i),f b (m))],
[0130] For target a in image A i And the target b that matches in image B. j Calculate the chi-square distribution values among them:
[0131]
[0132] Among them, v i [j] is the score vector v i The value of the j-th element is used to calculate d. i , indicating target a i The sum of the chi-square distribution values of the targets in all B images:
[0133]
[0134] If d i exceeds the threshold value d0, the match between the targets a i and b j is retained. Otherwise, the match is deleted. And according to the matching result, the tracking and management of the full life cycle of the multi-target sign in the panoramic image are performed, thereby improving the accuracy of data fusion and detection efficiency. An active traffic sign recognition algorithm based on the fusion of panoramic cameras and long-focus cameras is implemented, which combines the advantages of the panoramic camera with a large field of view and the long-focus camera with high-definition imaging at a long distance, thereby ensuring the completeness of detection and realizing accurate recognition of distant targets.
[0135] Based on the same inventive concept, the embodiments of the present specification also provide an active traffic sign recognition device based on the fusion of multiple cameras, as shown in Figure 7 , which comprises:
[0136] A preliminary recognition unit 701 is configured to collect panoramic images by a vehicle-mounted panoramic camera during vehicle driving and perform preliminary recognition to identify multiple signs in the panoramic images.
[0137] A confidence tracking unit 702 is configured to track the multiple signs with constantly changing relative positions due to vehicle driving, judge the confidence of the multiple signs detected preliminarily, and input the signs with a confidence lower than a preset threshold to a sorting module.
[0138] An importance sorting unit 703 comprises the sorting module and is configured to calculate an importance index of each sign with a confidence lower than a preset threshold, sort the signs according to the importance index, and determine the sign with the highest importance index as a current target sign.
[0139] A high-definition collection unit 704 is configured to adjust the target angle of a long-focus camera according to the changing relative position of the current target sign during vehicle driving, collect a high-definition image of the current target sign located at the target angle, sort the next sign as the current target sign according to the importance index, and sequentially collect high-definition images of the current target sign.
[0140] A high-definition recognition unit 705 is configured to perform high-definition recognition on the high-definition images, perform data fusion on the recognition results of the high-definition recognition and the preliminary detection results, and generate a traffic sign recognition result.
[0141] Since the principles of the above device for solving problems are similar to those of the above method, the implementation of the above device can be referred to the implementation of the above method, and the repeated parts will not be described again.
[0142] As Figure 8A structural diagram of a computer device of an embodiment of the present specification is shown. The apparatus in the embodiment of the present specification can be the computer device in the embodiment, which executes the method of the embodiment of the present specification. The computer device 802 can include one or more processing devices 804, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 802 can also include any storage resources 806 for storing any kind of information, such as code, settings, data, etc. Without limitation, for example, the storage resources 806 can include any one or combination of the following: any type of RAM, any type of ROM, flash memory devices, hard disks, optical disks, etc. More generally, any storage resource can store information using any technology. Further, any storage resource can provide volatile or non-volatile retention of information. Further, any storage resource can represent a fixed or removable component of the computer device 802. In one case, the computer device 802 can perform any operation of the associated instructions when the processing device 804 executes the associated instructions stored in any storage resource or combination of storage resources. The computer device 802 also includes one or more drive mechanisms 808, such as a hard disk drive mechanism, an optical disk drive mechanism, etc., for interacting with any storage resources.
[0143] The computer device 802 can also include an input / output module 810 (I / O) for receiving various inputs (via input devices 812) and for providing various outputs (via output devices 814). One particular output mechanism can include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input devices 812, and the output devices 814 can also not be included, just as a computer device in a network. The computer device 802 can also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the above-described components together.
[0144] The communication links 822 can be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication links 822 can include any combination of hardwired links, wireless links, routers, gateway functionality, name processors, etc., governed by any protocol or combination of protocols.
[0145] The embodiment of the present specification also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above method.
[0146] The embodiments of the present specification also provide a computer readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the method described above.
[0147] It should be understood that the size of the sequence number of each process described above does not mean the order of execution in various embodiments of the embodiments of the present specification. The execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present specification.
[0148] It should also be understood that in the embodiments of the present specification, the term "and / or" only describes the association relationship of the associated objects, which means that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the embodiments of the present specification generally represents that the front and rear associated objects are in an "or" relationship.
[0149] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the disclosed embodiments in the embodiments of the present specification can be realized by electronic hardware, computer software or combination of both. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been described generally in the above description. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present specification.
[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0151] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the embodiments of the device described above are merely schematic; for example, the division of the units is only a logical function division; there can be another division manner in actual implementation; for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electric, mechanical or in other forms.
[0152] In addition, each function unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated into a unit. The integrated unit can be implemented in the form of hardware, or in the form of a software function unit.
[0153] When the integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (such as a personal computer, a processor, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various other media that can store program codes.
[0154] The principles and implementation manners of the embodiments of the present application are described in the embodiments of the present application, and the above embodiment descriptions are only for helping to understand the method and the core idea of the embodiments of the present application; meanwhile, for those skilled in the art, according to the ideas of the embodiments of the present application, the specific implementation manners and application ranges can be changed; and in view of the above, the content of the present application should not be understood as limiting the embodiments of the present application.
Claims
1. An active traffic sign recognition method based on multi-camera fusion, characterized in that, The method includes: While the vehicle is in motion, the on-board panoramic camera captures panoramic images and performs preliminary identification, recognizing multiple markers in the panoramic images. Track the multiple signs whose relative positions change continuously due to the vehicle's movement, determine the confidence level of the initially identified multiple signs, and input the signs with confidence levels lower than a preset threshold into the sorting module; The sorting module calculates the importance index of each of the flags whose confidence level is lower than a preset threshold, sorts them according to the importance index, and determines the flag with the highest importance index as the current target flag. Specifically, the current driving task of the vehicle is obtained, and the importance index of the sign is calculated according to preset rules and the type of the sign; the target angle of the telephoto camera is adjusted according to the relative position of the current target sign as the vehicle moves, and a high-definition image of the current target sign located at the target angle is acquired; the next sign is selected as the current target sign according to the importance index, and high-definition images of the current target signs are acquired sequentially. The high-definition image is subjected to high-definition recognition, and the results of the high-definition recognition are fused with the results of the preliminary recognition to generate traffic sign recognition results.
2. The active traffic sign recognition method based on multi-camera fusion according to claim 1, characterized in that, Determining the confidence level of the initially identified multiple markers further includes: The traffic sign recognition result is directly generated for signs with a confidence level higher than a preset threshold.
3. The active traffic sign recognition method based on multi-camera fusion according to claim 2, characterized in that, The ranking module further includes the following steps in calculating the importance index of each flag with a confidence level below a preset threshold: Identify the type of the sign, which includes at least: warning sign, prohibition sign, instruction sign, or directional sign.
4. The active traffic sign recognition method based on multi-camera fusion according to claim 3, characterized in that, Adjusting the target angle of the telephoto camera based on the changing relative position of the current target marker during the vehicle's movement further includes: Determine the initial coordinates of the center point of the current target marker in the panoramic image; The angle deviation is determined based on the vehicle's driving status, and the target angle is generated based on the initial coordinates and the angle deviation. The telephoto camera is adjusted according to the target angle, and the focal length of the telephoto camera is adjusted according to the focus evaluation of the panoramic image.
5. The active traffic sign recognition method based on multi-camera fusion according to claim 4, characterized in that, Adjusting the focal length of the telephoto camera based on the focus evaluation of the panoramic image further includes: Obtain the grayscale value of the panoramic image, and calculate the focus evaluation function based on the grayscale value; When the focus evaluation function is greater than a preset threshold, the target focal length is calculated; The target focal length is used as the focal length of the telephoto camera.
6. The active traffic sign recognition method based on multi-camera fusion according to claim 5, characterized in that, Performing high-definition recognition on the high-definition image, and fusing the results of the high-definition recognition with the results of the preliminary recognition to generate traffic sign recognition results further includes: Determine the number of traffic signs in the result of the high-definition recognition; If the result of the high-definition recognition contains only one traffic sign, then its type and confidence level are taken as the recognition result of the traffic sign. If the high-definition recognition result contains multiple traffic signs, the high-definition recognition result is fused with the preliminary recognition result to generate a traffic sign recognition result.
7. The active traffic sign recognition method based on multi-camera fusion according to claim 6, characterized in that, The process of fusing the high-definition recognition results with the preliminary recognition results to generate traffic sign recognition results further includes: The preliminary identification results are saved in the panoramic image, and the center coordinates of multiple traffic signs in the panoramic image are calculated. A virtual viewfinder is generated based on the center coordinates and the panoramic image; Calculate the feature vectors of the high-definition image and the multiple traffic signs in the virtual viewfinder, and match the traffic signs in the panoramic image with the traffic signs in the high-definition image; The matching results, along with the type and confidence level of each traffic sign, are used as the recognition results for the traffic sign.
8. An active traffic sign recognition device based on multi-camera fusion, characterized in that, The device includes: The preliminary identification unit is used to collect panoramic images and perform preliminary identification on the vehicle-mounted panoramic camera while the vehicle is in motion, identifying multiple signs in the panoramic images. A confidence tracking unit is used to track the multiple signs whose relative positions change continuously due to the movement of the vehicle, determine the confidence level of the initially identified multiple signs, and input the signs with confidence levels lower than a preset threshold into the sorting module. The importance ranking unit includes the ranking module, which is used to calculate the importance index of each of the signs whose confidence level is lower than a preset threshold, rank them according to the importance index, and determine the sign with the highest importance index as the current target sign. Specifically, the current driving task of the vehicle is obtained, and the importance index of the sign is calculated according to preset rules and the type of the sign; The high-definition acquisition unit is used to adjust the target angle of the telephoto camera according to the relative position of the current target sign as the vehicle moves, and to acquire a high-definition image of the current target sign located at the target angle. The next sign is selected as the current target sign according to the importance index, and high-definition images of the current target signs are acquired sequentially. The high-definition recognition unit is used to perform high-definition recognition on the high-definition image, and to fuse the results of the high-definition recognition with the results of the preliminary recognition to generate traffic sign recognition results.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Traffic sign detection method and related equipment
CN113221756A
Traffic sign information extraction method and system based on image recognition
CN119418302A