An information matching method for traffic signs and related devices
By using an end-to-end neural network model in map applications, the text information in the target image is matched with traffic signs, and the problem of determining the matching relationship between multiple traffic signs and text information is solved, achieving efficient and accurate information matching.
Patent Information
- Application Number
- CN202110473889.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-29
AI Technical Summary
In map applications, how to accurately determine the matching relationship between text information and traffic signs, especially in the case where multiple traffic signs and multiple text information exist in the target image.
An end-to-end neural network is used as an information matching model to process the target image to determine the matching result between text information and traffic signs. This model improves the simplicity and robustness of the algorithm through the integration of feature extraction, detection and relationship matching.
It realizes accurate matching between text information and traffic signs, improves the efficiency and accuracy of information matching, and is suitable for scenarios such as automatic production of map data.
Smart Images

Figure CN113762039B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method for matching information of traffic signs and related devices. Background Art
[0002] In the process of map applications (including but not limited to scenarios such as navigation applications, autonomous driving, intelligent transportation, and map data production), it is necessary to accurately identify traffic signs. In some cases, traffic signs need to be combined with explanatory and supplementary texts (such as: affiliated information) to accurately and completely express their meanings. When performing identification, there are often multiple traffic signs and multiple text information in the acquired images. In this case, it is necessary to determine the matching relationship between the text information and the traffic signs, for example: which text information belongs to which signs, or which text information is combined with which signs to jointly express complete information.
[0003] How to accurately determine the matching relationship between text information and traffic signs is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] To solve the above technical problems, the present application provides a method for matching information of traffic signs and related devices.
[0005] In a first aspect, an embodiment of the present application provides a method for matching information of traffic signs, the method comprising:
[0006] Obtain a collected target image, where the target image includes text information and traffic signs;
[0007] Process the target image through an information matching model to obtain a matching result between the text information and the traffic signs in the target image, where the information matching model is an end-to-end neural network, and the neural network takes the target image as an input and the matching result between the text information and the traffic signs as an output.
[0008] In a second aspect, an embodiment of the present application provides a device for matching information of traffic signs, the device comprising an obtaining unit and a matching unit:
[0009] The obtaining unit is configured to obtain a collected target image, where the target image includes text information and traffic signs;
[0010] The matching unit is configured to process the target image through an information matching model to obtain a matching result between the text information and the traffic signs in the target image, where the information matching model is an end-to-end neural network, and the neural network takes the target image as an input and the matching result between the text information and the traffic signs as an output.
[0011] In a third aspect, an embodiment of the present application provides an information matching device for a traffic sign, and the device includes a processor and a memory:
[0012] The memory is used to store program code and transmit the program code to the processor;
[0013] The processor is used to execute the method described in the first aspect according to the instructions in the program code.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect.
[0015] It can be seen from the above technical solutions that when the embodiment of the present application performs information matching of traffic signs, it first obtains the collected target image, and the target image includes text information and traffic signs. Then, the target image is input into the information matching model, and through the information matching model, the target image is processed to obtain the matching result between the text information and the traffic signs in the target image. Among them, the information matching model is an end-to-end neural network with the target image as the input and the matching result between the text information and the traffic signs as the output. That is, the information matching model is a model obtained by overall training and optimization of the parameters of the entire neural network through an end-to-end training method. The overall detection performance of the information matching model is better. Therefore, the embodiment of the present application uses the information matching model to perform information matching of traffic signs, which can ensure the accuracy of the matching result and improve the efficiency of information matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is an example diagram of a traffic sign provided by an embodiment of the present application;
[0018] Figure 2 It is an architecture flowchart of information matching of a traffic sign provided by an embodiment of the present application;
[0019] Figure 3 It is a scenario architecture diagram of a method for information matching of a traffic sign provided by an embodiment of the present application;
[0020] Figure 4Flowchart of an information matching method for a traffic signboard provided by an embodiment of the present application;
[0021] Figure 5 Example diagram of a traffic signboard and text information provided by an embodiment of the present application;
[0022] Figure 6 Another example diagram of a traffic signboard and text information provided by an embodiment of the present application;
[0023] Figure 7 Example diagram of obtaining an affiliation graph based on a target image provided by an embodiment of the present application;
[0024] Figure 8 Schematic diagram of the architecture of an information matching model provided by an embodiment of the present application;
[0025] Figure 9 Schematic diagram of the architecture of a detection network provided by an embodiment of the present application;
[0026] Figure 10 Flowchart of information matching based on an information matching model provided by an embodiment of the present application;
[0027] Figure 11a Architecture flowchart of matching by a relationship matching network provided by an embodiment of the present application;
[0028] Figure 11b Example diagram of cropping a 2×2 feature map based on RoIAlign provided by an embodiment of the present application;
[0029] Figure 12 Flowchart of training an information matching model provided by an embodiment of the present application;
[0030] Figure 13 Structure diagram of an information matching device for a traffic signboard provided by an embodiment of the present application;
[0031] Figure 14 Structure diagram of a terminal provided by an embodiment of the present application;
[0032] Figure 15 Structure diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0033] Next, the embodiments of the present application will be described with reference to the accompanying drawings.
[0034] First, the nouns involved in the embodiments of the present application will be explained:
[0035] Traffic sign: A sign that can independently convey certain road indication information in a road scene. The abbreviations of "traffic sign" can be "sign", "label", etc. For example, Figure 1 As shown, traffic sign 1 can independently indicate that the speed limit for the current road section is "110 km / h", and traffic sign 2 can independently indicate that the speed limit for the current road section is "100 km / h".
[0036] Ancillary information: Text information located near the traffic sign and used to supplement and explain the traffic sign, such as information on time periods, road names, vehicle types, etc. For example, Figure 1 As shown, text information 1 is the ancillary information of traffic sign 1, and text information 1 further supplements and explains that "speed limit 110 km / h" only applies to "passenger cars with 7 seats or less". Text information 2 is the ancillary information of traffic sign 2, indicating that the speed limit for "other vehicle types" on this road section is "100 km / h".
[0037] Information matching of traffic signs: Matching and combining multiple traffic signs and multiple text information that appear in the target image. Multiple traffic signs and multiple text information often appear simultaneously in the target image. Information matching of traffic signs can be to determine whether each text information is ancillary information and which traffic sign's ancillary information each text information belongs to. For example, Figure 1 As shown, there are two traffic signs and two text information in this target image. Information matching of traffic signs needs to be able to determine that text information 1 is the ancillary information of traffic sign 1 and text information 2 is the ancillary information of traffic sign 2.
[0038] In one or more embodiments, in order to achieve information matching, two detectors are used to detect the text and traffic signs in the target image respectively. Then, the respectively obtained detection results (such as text detection results and traffic sign detection results) and the target image are sent to the matching model again to determine whether a certain text is the ancillary information of a traffic sign. See Figure 2 As shown, Figure 2 The gray box, dotted box, and black box in [Figure] respectively correspond to the method flows of the two detectors and the matching model, which are three models. As can be seen from Figure 2 [Figure], the two detectors and the matching model are three models independently trained. The three are trained with different training objectives respectively, rather than optimizing and adjusting the model parameters with a unified training objective, which is likely to cause error accumulation. In some scenarios, it will lead to a low accuracy rate of information matching. And in this process, each model needs to perform feature extraction once, resulting in the defect of low processing efficiency.
[0039] To solve the above technical problems, in one or more embodiments, text information detection, traffic sign detection, and the matching of traffic signs and text information are integrated into one model (i.e., the information matching model). By adopting an end-to-end information matching model, not only the simplicity and robustness of the algorithm are improved, but also the detection task and the matching task can promote and improve each other, thus realizing the joint optimization of the detection and matching parts. The embodiments of the present application use this information matching model to perform information matching of traffic signs, which can ensure the accuracy of the matching result and the efficiency of information matching.
[0040] The fields to which the method provided by the embodiments of the present application can be applied include, but are not limited to, fields such as maps, navigation, autonomous driving, and intelligent transportation. Taking the map field as an example, it can be used for automated map data production. By obtaining the matching result between text information and traffic signs through this method, the matching result can be provided for use in automated map data production services, so as to automatically produce more accurate road data and reduce the manual operation cost.
[0041] It should be noted that the method provided by the embodiments of the present application may involve artificial intelligence. Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results of theory, method, technology, and application system. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.
[0042] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0043] For example, computer vision technology in artificial intelligence software technology, Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to using cameras and computers to replace human eyes to identify targets, trace traces, measure targets, and perform other machine vision, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition. This application mainly relates to image semantic understanding technology, such as extracting image semantic features based on target images.
[0044] Another example is natural language processing technology in artificial intelligence. Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies. This application may, for example, involve semantic understanding in natural language processing, thereby detecting text information and obtaining text detection results.
[0045] Another example is machine learning in artificial intelligence. Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning. This application mainly trains information matching models through machine learning / deep learning.
[0046] See also Figure 3 , Figure 3Schematic diagram of an application scenario for the information matching method of traffic signs provided by an embodiment of this application. In this application scenario, a terminal 301 and a server 302 are included. The server 302 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal 301 may be an electronic device with an image acquisition function, such as a driving recorder, an in-vehicle camera, an in-vehicle computer, a smart phone, a tablet computer, a laptop computer, etc., but is not limited thereto. The terminal 301 and the server 302 may be directly or indirectly connected through wired or wireless communication means, which is not limited in this application.
[0047] Among them, the terminal 301 may be a device for collecting a target image. The terminal 301 may upload the collected target image to the server 302 so that the server 302 can execute the information matching method of traffic signs provided by an embodiment of this application. Among them, the terminal 301 may collect the target image through target image crowdsourcing, that is, outsource the collection task of the target image to various companies or individuals in society, and let them take and upload the target image of a specific location through the terminal 301.
[0048] After the server 302 obtains the target image including text information and traffic signs, it may process the target image through an information matching model to obtain a matching result between the text information in the target image and the traffic signs. For example Figure 3 as shown Figure 3 The text information in the target image shown includes text information 1 and text information 2, and the traffic signs include traffic sign 1 and traffic sign 2. Through this information matching model, the matching between the traffic signs and the text information can be realized to obtain a matching result, which indicates that text information 1 has an affiliated relationship with traffic sign 1, and text information 2 has an affiliated relationship with traffic sign 2, that is, text information 1 is the affiliated information of traffic sign 1, and text information 2 is the affiliated information of traffic sign 2.
[0049] Since this information matching model is an end-to-end neural network with the target image as the input and the matching result between the text information and the traffic signs as the output, that is, this information matching model is a model obtained by overall training and optimization of the parameters of the entire neural network through an end-to-end training method, and its overall detection performance is better, and the joint optimization of the detection and matching parts can be realized. Therefore, the embodiment of this application uses this information matching model to perform information matching of traffic signs, which can ensure the accuracy of the matching result and improve the efficiency of information matching.
[0050] Of course, in this embodiment, the acquired target image may also be sent to a specific terminal, and the method provided in the embodiments of the present application may be executed by the terminal, or the method provided in the embodiments of the present application may also be executed in cooperation with the terminal and the server. This embodiment does not make any limitation in this regard.
[0051] Next, taking the server as the execution subject as an example, the network address conversion method provided in the embodiments of the present application will be introduced in detail with reference to the accompanying drawings.
[0052] See Figure 4 , Figure 4 which shows a flowchart of an information matching method for traffic signs, and the method includes:
[0053] S401. Obtain the acquired target image.
[0054] Among them, the target image can be realized through target image crowdsourcing, that is, outsourcing the acquisition task of the target image to various companies or individuals in society. Various companies or individuals in society can take pictures of roads at specific locations through terminals with image acquisition functions to obtain the target image, and then upload the acquired target image to the server.
[0055] The target image includes text information and traffic signs. Since the target image may include multiple traffic signs and multiple text information, it is necessary to determine which text information is the affiliated information of which traffic sign to obtain a matching result for subsequent use of the matching result, such as applying it to the automated production of map data.
[0056] Generally, a traffic sign may appear together with its corresponding affiliated information, or may not have corresponding affiliated information (that is, the traffic sign does not require text information to supplement it), and of course, only text information may appear (that is, the text information is not used as the affiliated information of any traffic sign). In addition, this embodiment does not make any limitation on the positional relationship between the text information and the traffic sign. The text information may be located below, to the left, to the right, above, etc. of the traffic sign.
[0057] See Figure 5 As shown, the text information "car" is on the left side of the traffic sign (the traffic sign can be seen as shown by the dashed box 501 in Figure 5 ), and the text information "large vehicle" is on the left side of the traffic sign (the traffic sign can be seen as shown by the dashed box 502 in Figure 5 ); while Figure 1 in Figure 5 the text information 1 is below the traffic sign 1, and the text information 2 is below the traffic sign 2; in addition, in Figure 5 the text information "large vehicle keep right" exists alone and is not used as the affiliated information of the traffic sign shown in
[0058] See Figure 6 As shown, above the text information "trucks over 1.5t" there are two traffic signs, "No Entry" (see Figure 6 the dashed box in Figure 6 ), and "Speed Limit 20" (see
[0059] S402. Process the target image through an information matching model to obtain the matching result between the text information and traffic signs in the target image.
[0060] The server inputs the acquired target image into a pre-trained information matching model, and the information matching model processes the target image to obtain the matching result between the text information and traffic signs in the target image. This matching result can reflect which text information in the target image is the attached information of which traffic sign.
[0061] It should be noted that the information matching model used in this embodiment is an end-to-end neural network, which takes the target image as the input and the matching result between the text information and traffic signs as the output. That is, after the server inputs the acquired target image into this information matching model, the information matching model can correspondingly output the matching result.
[0062] It can be understood that the matching result can be represented in various forms. One way can be to directly represent in words which text information is the attached information of which traffic sign, such as "Text Information 1 is the attached information of Traffic Sign 1", "Text Information 2 is the attached information of Traffic Sign 2". Another way can be to represent the matching result through an attachment relationship diagram. See Figure 7 as shown Figure 7 In the left figure, it shows that the detected traffic signs in the target image include Traffic Sign 1, Traffic Sign 2, and Traffic Sign 3, and the text information includes Text Information 1, Text Information 2, and Text Information 3. The attachment relationship diagram obtained through the information matching model is as shown in Figure 7As shown in the middle right figure. In the affiliation relationship diagram, the element in the \(i\)-th row and \(j\)-th column indicates whether the text information \(i\) is the affiliated information of the traffic sign \(j\). If so, it is represented by 1; if not, it is represented by 0. Therefore, according to this affiliation relationship diagram, it can be determined that text information 1 is the affiliated information of traffic sign 1, and text information 2 is the affiliated information of traffic sign 2. However, text information 3 is not the affiliated information of any traffic sign, and traffic sign 3 does not have any text information as its affiliated information. Then, false detections may occur for text information 3 and traffic sign 3. However, since the information matching model does not match these two possible false detection results together, it avoids information matching errors caused by false detections.
[0063] It should be noted that in this embodiment, referring to Figure 7 the example shown, the target image includes multiple traffic signs and multiple text information. These multiple traffic signs and multiple text information form multiple pairs of traffic signs and text information whose affiliation relationship needs to be determined. The information matching model used in the embodiments of the present application can perform reasoning and judgment on all possible affiliation relationships existing in the target image at one time, which can greatly improve the efficiency of information matching.
[0064] In the embodiments of the present application, after obtaining the matching result, the text information and traffic signs with an affiliation relationship can be provided for use in the automated production service of map data, so as to automatically produce more accurate road data and reduce the manual operation cost.
[0065] As can be seen from the above technical solution, when performing information matching of traffic signs in the embodiments of the present application, first, the acquired target image is obtained, and the target image includes text information and traffic signs. Then, the target image is input into the information matching model, and through the information matching model, the target image is processed to obtain the matching result between the text information and traffic signs in the target image. Among them, the information matching model is an end-to-end neural network that takes the target image as the input and the matching result between the text information and traffic signs as the output. That is, the information matching model is a model obtained by overall training and optimization of the parameters of the entire neural network through an end-to-end training method. The overall detection performance of the information matching model is better. Therefore, using the information matching model in the embodiments of the present application for information matching of traffic signs can ensure the accuracy of the matching result and improve the efficiency of information matching.
[0066] As described above, the information matching method for traffic signs provided by the embodiments of the present application needs to be based on an information matching model to determine the matching result according to the acquired target image. To facilitate further understanding of the specific implementation process of the information matching method for traffic signs provided by the embodiments of the present application, the above information matching model will be specifically introduced below with reference to the accompanying drawings.
[0067] See Figure 8 , Figure 8 , which is a schematic diagram of the architecture of the information matching model provided by the embodiment of the present application. As Figure 8 shown, the information matching model includes a feature extraction network 801, a detection network 802, and a relationship matching network 803.
[0068] Among them, the feature extraction network 801 is a first neural network that takes a target image as input and outputs image semantic features.
[0069] As the first neural network in the information matching model, the feature extraction network 801 is responsible for extracting features from the target image input to the information matching model to obtain image semantic features, and outputting the extracted image semantic features to the second neural network in the information matching model.
[0070] The feature extraction network 801 uses a deep convolutional neural network to extract rich features of the target image at different scales and different semantic dimensions. The feature extraction network 801 may include a convolutional layer, a pooling layer, a normalization layer, and an activation layer. The convolutional layer uses convolutional kernels of different sizes, such as 3×3, 5×5, 7×7, etc., to calculate each position in the target image in turn to extract basic texture features in the target image. The pooling layer can reduce the resolution of the features through pooling operations and map low-level semantic features to high-level semantic features. The normalization layer normalizes the output of each convolutional layer, so that the output of each layer can satisfy the normal distribution, which can accelerate the convergence of the model and improve the accuracy of the model at the same time. The activation layer performs a non-linear mapping on the features, changing the simple features to a higher-dimensional space to improve the expression ability of the model. The target image passes through the convolutional layer, pooling layer, normalization layer, and activation layer in the feature extraction network 801 in turn to obtain the final image semantic features.
[0071] The dimension of the image semantic features finally output by the feature extraction network 801 (as shown by the white square in Figure 9 ) is (N, C, W, H), where N is the number of target images processed in each batch, C is the number of feature channels, and W and H are the width and height of the feature map corresponding to the image semantic features, respectively.
[0072] It can be understood that in the embodiments of the present application, convolutional neural networks with other structures can also be used to extract image semantic features, and the structure of the above-mentioned feature extraction network 801 is only an example. In some possible implementation manners, in order to improve the accuracy of information matching, the accuracy of feature extraction can be improved first. Therefore, the embodiments of the present application further improve the feature extraction network 801, and design a more complex backbone network for the feature extraction network 801, such as designing more complex backbone networks such as res2net and hrnet, so as to improve the accuracy of feature extraction when using the feature extraction network 801 to extract image semantic features, and further improve the accuracy of information matching.
[0073] In other possible implementation manners, in order to improve the efficiency of information matching, the efficiency of feature extraction can be improved first. Therefore, the embodiments of the present application further improve the feature extraction network 801, and design a more lightweight backbone network for the feature extraction network 801 to accelerate the feature extraction efficiency, and further improve the information matching efficiency. For example, lightweight backbone networks such as moblienet and sufflenet can be designed to accelerate the algorithm running speed. The structure of the convolutional neural network as the feature extraction network 801 is not limited in any way here.
[0074] The detection network 802 is a second neural network that takes the output of the feature extraction network 801 as input and outputs text detection results and traffic sign detection results. The text detection results include text position information indicating the position of the text information in the target image, and the traffic sign detection results include sign position information indicating the position of the traffic sign in the target image.
[0075] That is to say, the detection network 802 is the second neural network in the information matching model, and is responsible for detecting according to the image semantic features output by the feature extraction network 801, and determining the text detection results and traffic sign detection results, that is, determining the traffic signs included in the target image and their corresponding positions, as well as the text information included and their corresponding positions.
[0076] In some cases, due to the large differences in the appearance forms and semantic information between traffic signs and text information, using the same detection network to detect these two types of targets at the same time may have poor effects. Therefore, in order to improve the detection accuracy, the embodiments of the present application design two different branches to complete the detection task. As Figure 9 shown, at this time, the detection network 802 includes a text information detection branch 8021 and a traffic sign detection branch 8022.
[0077] Among them, the text information detection branch 8021 is a fourth neural network that takes the output of the feature extraction network 801 as input and the text detection result as output. The text information detection branch 8021 is the fourth neural network in the detection network 802 and is responsible for determining the text detection result according to the image semantic features output by the feature extraction network 801.
[0078] The traffic sign detection branch 8022 is a fifth neural network that takes the output of the feature extraction network 801 as input and the traffic sign detection result as output. The traffic sign detection branch 8022 is the fifth neural network in the detection network 802 and is responsible for determining the traffic sign detection result according to the image semantic features output by the feature extraction network 801.
[0079] The relationship matching network 803 is a third neural network that takes the output of the detection network 802 as input and the matching result between the text information and the traffic signs in the target image as output.
[0080] As the third neural network in the information matching model, the relationship matching network 803 is responsible for matching the text detection result and the traffic sign detection result output by the detection network 802 to obtain the matching result, so as to determine which text information in the target image is the attached information of which traffic sign.
[0081] In some cases, the relationship matching network 803 also takes the output of the feature extraction network 801 as input, that is, the image semantic features output by the feature extraction network 801 can be input into the relationship matching network 803, so that the relationship matching network 803 can determine the matching result according to the image semantic features, the text detection result and the traffic sign detection result at the same time.
[0082] The above information matching model includes a feature extraction network, a detection network and a relationship matching network. Correspondingly, when using this information matching model for information matching, according to the input target image, the matching result can be determined in one step through the feature extraction network, the detection network and the relationship matching network included in the information matching model.
[0083] Based on Figure 8 the information matching model shown, when performing information matching, Figure 4 the specific implementation of the information matching method for the traffic signs shown can refer to Figure 10 , Figure 10 For Figure 8 the flowchart of information matching based on the information matching model shown, the method includes:
[0084] S1001. Obtain the target image.
[0085] S1001 and Figure 4Similar to the specific implementation of S401 in the corresponding embodiment, details are not described here. Refer to the relevant description of S401.
[0086] S1002. Extract image semantic features from the target image through the feature extraction network.
[0087] After obtaining the target image, input the target image into the feature extraction network of the information matching model. The feature extraction network uses the convolutional neural network model included therein to extract image semantic features from the target image. Then, input the obtained image semantic features into the detection network of the information matching model.
[0088] S1003. Determine the text detection result and the traffic sign detection result through the detection network according to the image semantic features.
[0089] The detection network determines the text detection result and the traffic sign detection result according to the input image semantic features, and then inputs the text detection result and the traffic sign detection result into the relationship matching network.
[0090] It should be noted that in some cases, in order to visualize the text detection result and the traffic sign detection result, the detection result can be further parsed into a detection box. For example, the text detection result can be represented by the first detection box in the target image, such as Figure 9 shown by the dashed rectangular box in; represent the traffic sign detection result by the second detection box in the target image, such as Figure 9 shown by the gray solid rectangular box in.
[0091] If the detection network is as shown in 9 and includes a text information detection branch and a traffic sign detection branch, since Figure 9 the two different branches in are respectively used to detect different targets (text information or traffic signs), and the appearance forms and semantic information of these two different targets are quite different, in order to accurately detect which are text information and which are traffic signs, the features of the detected targets can be made more prominent, that is, make the features of the text information more prominent when detecting text information, and make the features of the traffic signs more prominent when detecting traffic signs.
[0092] In a possible implementation, the implementation of S1003 can be to perform feature transformation on the image semantic features through the text information detection branch to obtain the first semantic feature. The significance of the features of the text information in the first semantic feature is higher than that of the features of the text information in the image semantic features, thereby making the features of the text information more prominent. Then, based on the first semantic feature, the text detection result is obtained through the text information detection branch. Since the features of the text information are more prominent at this time, it is beneficial to the detection of the text information, making the obtained text detection result more accurate. The traffic sign detection branch performs feature transformation on the image semantic features to obtain the second semantic feature. The significance of the features of the traffic signs in the second semantic feature is higher than that of the features of the traffic signs in the image semantic features, thereby making the features of the traffic signs more prominent. Then, based on the second semantic feature, the traffic sign detection result is obtained through the traffic sign detection branch. Since the features of the traffic signs are more prominent at this time, it is beneficial to the detection of the traffic signs, making the obtained traffic sign detection result more accurate.
[0093] See Figure 9 as shown in Figure 9 The white cube in represents the image semantic features output by the feature extraction network, and this image semantic feature is shared by the traffic sign detection branch and the text information detection branch. Taking the traffic sign detection branch 8022 as an example, this traffic sign detection branch will perform further feature extraction and feature transformation on the image semantic features (white cube), making the features of the traffic signs more prominent, so as to obtain the second semantic feature (light gray cube) that is more beneficial to traffic sign detection. After that, a 1×1 convolution is used to transform the feature dimension and output the traffic sign detection result (the gray solid rectangle frame in the upper right), and the included sign position information represents the distance from this point to the four bounding boxes of up, down, left, and right. Similarly, the text information detection branch 8021 obtains the first semantic feature (dark gray cube) based on the image semantic features, making the features of the text information more prominent, so as to accurately predict the position of the text information in the target image and obtain the text detection result.
[0094] In some cases, in order to represent the credibility of the detection results, so that the detection results with higher credibility can be selected for subsequent matching to improve the matching efficiency in the future, the text detection result also includes the first confidence score corresponding to the text detection result, and the traffic sign detection result also includes the second confidence score corresponding to the traffic sign detection result. The first confidence score is used to represent the credibility of the text detection result, and the second confidence score is used to represent the credibility of the traffic sign detection result, so as to exclude some untrustworthy detection results before matching and improve the matching efficiency.
[0095] At this time, the number of channels corresponding to the text detection result is 5, which respectively represent the distances from this point to the four bounding boxes of up, down, left, and right and the first confidence score. The number of channels corresponding to the traffic sign detection result is 5, which respectively represent the distances from this point to the four bounding boxes of up, down, left, and right and the second confidence score.
[0096] The following introduces the detection process of the detection network in combination with the specific network structures of the text information detection branch and the traffic sign detection branch. The network structure of the traffic sign detection branch consists of four 3×3 convolutions and one 1×1 convolution. The 4 3×3 convolutions perform further non-linear transformations on the shared image semantic features, making the features of the traffic signs more significant and prominent. The 1×1 convolution serves as the output layer, reducing the channel dimension to 5 channels, and the final output dimension is (N, 5, W, H). The 5 channels here respectively represent the distances from this point to the four bounding boxes of up, down, left, and right and the second confidence score.
[0097] The network structure of the text information detection branch is similar to that of the traffic sign detection branch, including 3×3 convolutions for enhancing text features and 1×1 convolutions for dimension transformation. In addition, compared with the enumerable styles of traffic signs, the styles of text information vary greatly. Therefore, after the 4 3×3 ordinary convolutions, 4 3×3 deformable convolutions are added to its network structure. The deformable convolution can determine the deformation parameters of the convolution based on learning, expanding the receptive field of the convolution, so that the feature extraction ability of the text information detection branch is more powerful, thus ensuring the effect of text information detection.
[0098] In addition, due to the different shape distributions of traffic signs and text information, the length and width of traffic signs are mostly close, while text information is mostly in the shape of "long strips". Therefore, during training, in order to make both detection branches achieve the optimal detection effect, different prior boxes can be used in the two detection branches (prior boxes are generally used to preset the length and height of the target such as traffic signs or text information to assist in prediction), and the specific shape of the prior box can be obtained by clustering on the training data. Through such a design, not only can the convergence of the model be accelerated, but also both detection branches can be fully trained to achieve the optimal training effect at the same time.
[0099] S1004. According to the text position information in the text detection result and the sign position information in the traffic sign detection result, determine the matching result between the text information and the traffic sign in the target image through the relationship matching network.
[0100] After inputting the text detection result and the traffic sign detection result into the relationship matching network, the relationship matching network can determine the matching result between the text information and the traffic sign in the target image according to the text position information in the text detection result and the sign position information in the traffic sign detection result.
[0101] In some cases, if the relationship matching network also takes the output of the feature extraction network as input, then when determining the matching result, the position information and the image semantic features can be comprehensively used as the inference and judgment basis for the affiliated relationship. Specifically, a possible implementation manner of determining the matching result may be to determine the sign position encoding corresponding to the sign position information and the text position encoding corresponding to the text position information, and then fuse the sign position encoding, the text position encoding, and the image semantic features to obtain a fused feature, so as to determine the matching result between the text information and the traffic sign in the target image according to the fused feature.
[0102] See Figure 11a As shown, the input of the relationship matching network is divided into two parts. One part is the text detection result and the traffic sign detection result output by the detection network, and the other part is the image semantic features output by the feature extraction network. The relationship matching network first performs Position Embedding, that is, position encoding, on the sign position information and the text position information, so as to better extract the spatial position information between different detection frames, and at the same time, it can also be more convenient to fuse with the image semantic features. Among them, both the sign position information and the text position information can be represented in the coordinate form of [x1, y1, x2, y2].
[0103] The specific method of position encoding is as follows: First, generate a two-dimensional array with the same scale as the target image, and its initial value is all 0. For the sign position information corresponding to a certain traffic sign or the text position information corresponding to the text information [x1, y1, x2, y2], set the values of the points within the rectangular area enclosed by these two points (x1, y1) to (x2, y2) to 1. In this way, a 0-1 binary image that can represent the position of this traffic sign or text information is obtained as the sign position encoding or the text position encoding. For how many sign position information and text position information need to be position-encoded, that many binary images can be obtained, and the obtained binary images are stacked along the channel dimension. Finally, it is scaled to a specific scale through bilinear interpolation to facilitate fusion with the image semantic features. Among them, the specific scale can be the same as the scale of the image semantic features, such as a scale of 7×7. In this embodiment, the final output dimension of the position encoding can be K×1×7×7, and K can represent the number of detection frames (including the first detection frame or the second detection frame) that need to be matched.
[0104] This application makes full use of location information and image semantic features, and comprehensively applies spatial constraints and semantic constraints during the information matching process to effectively improve the accuracy of information matching.
[0105] In some cases, since the image semantic features output by the feature extraction network are features of the entire image, these image semantic features may include a lot of irrelevant or even interfering background features. Therefore, based on the text detection results and traffic sign detection results output by the detection network, embodiments of this application crop out the features of the region of interest (RoI). To this end, in embodiments of this application, a possible implementation of obtaining the fused features may be to determine the region of interest according to the sign position information and text position information, crop out the corresponding semantic features of the region of interest from the image semantic features, and then fuse the semantic features of the region of interest with the corresponding sign position encoding or text position encoding to obtain the fused features. The semantic features of the region of interest may include the semantic features of the region of interest corresponding to the traffic sign and the semantic features of the region of interest corresponding to the text information. That is to say, the semantic features of the region of interest cropped according to the sign position information (i.e., the semantic features of the region of interest corresponding to the traffic sign) are fused with the corresponding sign position encoding, and the semantic features of the region of interest cropped according to the text position information (i.e., the semantic features of the region of interest corresponding to the text information) are fused with the corresponding text position encoding, so as to obtain the fused features. Wherein, if the text detection result is represented by the first detection box and the traffic sign detection result is represented by the second detection box, the region of interest may be the region enclosed by the first detection box and the second detection box.
[0106] Specifically, embodiments of this application can complete this "cropping" operation based on RoIAlign. RoIAlign can uniformly map features of different sizes to a fixed dimension, such as 7×7, for subsequent inference of the affiliation relationship.
[0107] The following takes the example of cropping out the semantic features of the region of interest of 2×2 using RoIAlign to discuss its specific operation:
[0108] As Figure 11b shown, the black solid-line rectangular box represents the position of the target (such as a traffic sign or text information) on the feature map corresponding to the image semantic features. If it is desired to crop out the semantic features of the region of interest from the feature map corresponding to the image semantic features to obtain the feature map corresponding to the semantic features of the region of interest of 2×2 size, then the black solid-line rectangular box can be equally divided into four regions 1, 2, 3, and 4. Then, a value is calculated for each region to represent the region, and these four values form the required 2×2 feature map. Among them, the value of each region can be four sampling points evenly distributed within the region (such as Figure 11bThe average value of the black dots in
[0109] Then, the semantic features of interest obtained by RoIAlign cropping are stacked with the signboard position features or text position features to fully fuse the two features with different dimensions, resulting in fused features. Thus, in the following, it is possible to adaptively learn the contribution of each dimension feature to the final decision based on a large amount of data, and finally output an affiliation graph.
[0110] It should be noted that, in order to obtain the fused features, the relationship matching network provided by the embodiments of the present application may include two 1×1 convolutions, two 3×3 convolutions, a global pooling layer, and two fully connected layers.
[0111] Specifically, the two features for fusion can be the previously obtained position encoding of K×1×7×7 dimensions and the image semantic features (such as the semantic features of interest), and the dimension of the image semantic features can be K×C×7×7 dimensions. Since the channel number dimension of the position encoding is only 1 dimension, which is relatively small compared to the dimension of the image semantic features, it is easy for the position information represented by the position encoding to be overwhelmed by the image semantic features. At the same time, the previous position encoding only considered the absolute position of the target in the target image and did not consider the relative position between targets. Therefore, the embodiments of the present application use two 1×1 convolutions in the relationship matching network to first increase the channel number of the position encoding from 1 dimension to C / 2 dimensions and then to C dimensions, so as to obtain features that encode both the absolute position and the relative position, and the dimension also becomes K×C×7×7.
[0112] Then, two 3×3 convolutions are respectively used to further transform the position encoding of K×C×7×7 dimensions and the image semantic features of K×C×7×7 dimensions. The two transformed features are then stacked together to become K×2C×7×7 dimensions. After that, 3×3 convolutions and 1×1 convolutions are alternately used to continue performing non-linear transformation and information fusion on the stacked features. Then, a global pooling layer is used to transform the K×2C×7×7 features into K×2C×1×1, that is, K×2C dimensions. Then, two fully connected layers are used to continue fusing the image semantic features and the position encoding, and finally, fused features of K×K dimensions are output.
[0113] It should be noted that if the text detection result also includes the first confidence score corresponding to the text detection result, and the traffic sign detection result also includes the second confidence score corresponding to the traffic sign detection result, in order to avoid matching some misdetection results, reduce the computation amount, and improve the matching efficiency, a possible implementation manner of S1004 may be to select M text detection results according to the first confidence score, and select N traffic sign detection results according to the second confidence score and input them into the affiliation matching network. The affiliation matching network outputs an affiliation graph, which is used to represent the matching result between the text information corresponding to the M text detection results and the traffic signs corresponding to the N traffic sign detection results. Among them, M and N may be equal or not equal, and this embodiment does not limit this. Figure 7 Taking M = N = 3 as an example for introduction.
[0114] Among them, the selection method may include various types. One method may be to sort the first confidence scores in descending order and select the first M text detection results, and sort the second confidence scores in descending order and select the first N traffic sign detection results. Another method may be to sort the first confidence scores in ascending order and select the last M text detection results, and sort the second confidence scores in ascending order and select the last N traffic sign detection results. In some cases, it may also be to select M text detection results with the first confidence score greater than the first threshold and select N traffic sign detection results with the second confidence score greater than the second threshold. This embodiment does not limit the selection method.
[0115] In this embodiment, the detection results with higher confidence scores are selected from all the detection results for continued matching, so that it is not necessary to calculate each detection result, reducing the computation amount and improving the matching efficiency.
[0116] It can be understood that whether the above information matching model can accurately determine the matching result depends on the model performance of the information matching model, and the quality of the model performance of the information matching model depends on the training process of the information matching model.
[0117] Next, in combination with Figure 12 the process of training the information matching model will be introduced. Refer to Figure 12 , the method includes:
[0118] S1201. Construct an initial information matching model, where the initial information matching model includes an initial feature extraction network, an initial detection network, and an initial relationship matching network.
[0119] Taking the constructed initial information matching model as the training basis, train the initial information matching model. It can be understood that the structure of the initial information matching model is similar to that of the information matching model, including an initial feature extraction network, an initial detection network, and an initial relationship matching network.
[0120] S1202. Obtain the training samples in the training sample set. The training samples include training images and the true matching results between the text information and the traffic signs.
[0121] When training the initial information matching model, it is necessary to obtain the training samples in the training sample set and use these training samples to train the constructed initial information matching model.
[0122] Since the input of the information matching model is the target image and the output is the matching result, when training the initial information matching model with the training samples, it is necessary to obtain the same input and output as the information matching model, that is, the training samples obtained need to include the training images and the true matching results between the text information and the traffic signs, so as to ensure that the information matching model trained with these training samples can meet the input and output requirements of the information matching model in practical applications.
[0123] S1203. Input the training image into the initial information matching model, and successively process it through the initial feature extraction network, the initial detection network, and the initial relationship matching network to obtain the output content of the initial relationship matching network. The output content includes the predicted matching result between the text information and the traffic sign.
[0124] Input the training image into the initial information matching model. Use the initial feature extraction network in the initial information matching model to extract the image semantic features, and input the image semantic features into the initial detection network. The initial detection network predicts the text detection result and the traffic sign detection result according to the image semantic features, and inputs the text detection result and the traffic sign detection result into the initial relationship matching network. The initial relationship matching network performs matching according to the text detection result and the traffic sign detection result to obtain the predicted matching result.
[0125] S1204. Construct a loss function according to the predicted matching result and the true matching result.
[0126] S1205. Adjust the model parameters of the initial information matching model according to the loss function. According to the adjusted model parameters and the network structure of the initial information matching model when the training conditions are met, determine the information matching model.
[0127] A loss function is constructed based on the error between the predicted matching result and the true matching result output by the initial information matching model. Furthermore, based on this loss function, the model parameters in the initial information matching model can be adjusted, thereby realizing the optimization of the initial information matching model. When the initial information matching model meets the training conditions, the information matching model can be determined according to the model parameters of the current initial information matching model and the network structure of the initial information matching model.
[0128] When adjusting the model parameters, the loss represented by the loss function can be backpropagated in the initial feature extraction network, the initial detection network, and the initial relationship matching network. According to the loss, the weights and model parameters of the initial feature extraction network, the initial detection network, and the initial relationship matching network are adjusted until the training conditions are met, so as to obtain the information matching model, thereby realizing the joint optimization of the initial detection network and the initial relationship matching network.
[0129] This application integrates the two tasks of traffic sign - text information detection and information matching, reducing the algorithm complexity through feature reuse. And through joint optimization, the detection, matching efficiency and matching accuracy of traffic sign - text information are greatly improved, the cost of map automated production can be effectively reduced, and the production quality can be significantly improved.
[0130] Next, the information matching method for traffic signs provided in the embodiments of this application will be introduced in combination with an actual application scenario. In this scenario, the target image can be collected through the vehicle's driving recorder and the collected target image is uploaded to the server. After the server obtains the target image, the target image can be input into the information matching model. The image semantic features are extracted from the target image through the feature extraction network in the information matching model, and the image semantic features are input into the detection network. The traffic sign detection branch in the detection network determines the traffic sign detection result according to the image semantic features and inputs the traffic sign detection result into the relationship matching network; the text information detection branch in the detection network determines the text detection result according to the image semantic features and inputs the text detection result into the relationship matching network. At the same time, the feature extraction network inputs the image semantic features into the relationship matching network, and the relationship matching network determines the matching result according to the image semantic features, the text position information in the text detection result, and the sign position information in the traffic sign detection result. This matching result can reflect which text information is the attached information of which traffic sign, and then the text information and traffic signs with an attached relationship are provided for the map data automated production service to use.
[0131] Based on Figure 4 corresponding to the information matching method for traffic signs provided in the embodiments, the embodiments of this application also provide an information matching device for traffic signs, see Figure 13, the device 1300 includes an acquisition unit 1301 and a matching unit 1302:
[0132] The acquisition unit 1301 is configured to acquire a target image obtained by collection, where the target image includes text information and a traffic sign;
[0133] The matching unit 1302 is configured to process the target image through an information matching model to obtain a matching result between the text information and the traffic sign in the target image. The information matching model is an end-to-end neural network that takes the target image as an input and the matching result between the text information and the traffic sign as an output.
[0134] In a possible implementation, the information matching model includes a feature extraction network, a detection network, and a relationship matching network;
[0135] Among them, the feature extraction network is a first neural network that takes the target image as an input and the image semantic feature as an output;
[0136] The detection network is a second neural network that takes the output of the feature extraction network as an input and the text detection result and the traffic sign detection result as outputs. The text detection result includes text position information indicating the position of the text information in the target image, and the traffic sign detection result includes sign position information indicating the position of the traffic sign in the target image;
[0137] The relationship matching network is a third neural network that takes the output of the detection network as an input and the matching result between the text information and the traffic sign in the target image as an output.
[0138] In a possible implementation, the matching unit 1302 is configured to:
[0139] Extract image semantic features from the target image through the feature extraction network;
[0140] Determine the text detection result and the traffic sign detection result through the detection network according to the image semantic feature;
[0141] Determine the matching result between the text information and the traffic sign in the target image through the affiliation relationship matching network according to the text position information in the text detection result and the sign position information in the traffic sign detection result.
[0142] In a possible implementation, if the affiliation relationship matching network also takes the output of the feature extraction network as an input, the matching unit 1302 is configured to:
[0143] Determine the sign position code corresponding to the sign position information and the text position code corresponding to the text position information;
[0144] Fuse the sign position code, the text position code, and the image semantic features to obtain fused features;
[0145] Determine the matching result between the text information and the traffic sign in the target image according to the fused features.
[0146] In a possible implementation, the matching unit 1302 is configured to:
[0147] Determine the region of interest according to the sign position information and the text position information;
[0148] Crop the region of interest corresponding to the region of interest from the image semantic features to obtain the region of interest semantic features;
[0149] Fuse the region of interest semantic features with the corresponding sign position code or text position code to obtain the fused features.
[0150] In a possible implementation, the detection network includes a text information detection branch and a traffic sign detection branch;
[0151] Among them, the text information detection branch is a fourth neural network with the output of the feature extraction network as the input and the text detection result as the output;
[0152] The traffic sign detection branch is a fifth neural network with the output of the feature extraction network as the input and the traffic sign detection result as the output.
[0153] In a possible implementation, the matching unit 1302 is configured to:
[0154] Perform feature transformation on the image semantic features through the text information detection branch to obtain the first semantic features, and the significance of the features of the text information in the first semantic features is higher than the significance of the features of the text information in the image semantic features;
[0155] According to the first semantic features, obtain the text detection result through the text information detection branch;
[0156] Perform feature transformation on the image semantic features through the traffic sign detection branch to obtain the second semantic features, and the significance of the features of the traffic sign in the second semantic features is higher than the significance of the features of the traffic sign in the image semantic features;
[0157] According to the second semantic feature, the traffic sign detection result is obtained through the traffic sign detection branch.
[0158] In a possible implementation manner, the text detection result is represented by a first detection box in the target image, and the traffic sign detection result is represented by a second detection box in the target image.
[0159] In a possible implementation manner, the text detection result further includes a first confidence score corresponding to the text detection result, and the traffic sign detection result further includes a second confidence score corresponding to the traffic sign detection result. The matching unit 1302 is configured to:
[0160] Select M text detection results according to the first confidence score, and select N traffic sign detection results according to the second confidence score and input them into the relationship matching network;
[0161] Output an affiliation relationship graph through the relationship matching network, where the affiliation relationship graph is used to represent the matching result between the text information corresponding to the M text detection results and the traffic signs corresponding to the N traffic sign detection results.
[0162] In a possible implementation manner, the device further includes a training unit:
[0163] The training unit is configured to:
[0164] Construct an information initial matching model, where the information initial matching model includes an initial feature extraction network, an initial detection network, and an initial relationship matching network;
[0165] Obtain training samples in the training sample set, where the training samples include training images and the true matching results between text information and traffic signs;
[0166] Input the training images into the information initial matching model, and after being processed by the initial feature extraction network, the initial detection network, and the initial relationship matching network in sequence, obtain the output content of the initial relationship matching network, where the output content includes the predicted matching result between the text information and the traffic signs;
[0167] Construct a loss function according to the predicted matching result and the true matching result;
[0168] Adjust the model parameters of the information initial matching model according to the loss function, and determine the information matching model according to the adjusted model parameters and the network structure of the information initial matching model when the training conditions are met.
[0169] In a possible implementation, the training unit is configured to:
[0170] Backpropagate the loss represented by the loss function in the initial feature extraction network, the initial detection network, and the initial relationship matching network, and adjust the weights and model parameters of the initial feature extraction network, the initial detection network, and the initial relationship matching network according to the loss until the training condition is met, thereby obtaining the information matching model.
[0171] An embodiment of the present application further provides an information matching device for traffic signs. This device may be a terminal. Taking a smartphone as an example of the terminal:
[0172] Figure 14 The block diagram of a part of the structure of the smartphone related to the terminal provided by the embodiment of the present application is shown. Refer to Figure 14 , the smartphone includes: a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490, etc. The input unit 1430 may include a touch panel 1431 and other input devices 1432. The display unit 1440 may include a display panel 1441. The audio circuit 1460 may include a speaker 1461 and a microphone 1462. Those skilled in the art can understand that Figure 14 the structure of the smartphone shown in
[0173] does not limit the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0174] The processor 1480 is the control center of the smart phone, connecting various parts of the entire smart phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1420, and by calling data stored in the memory 1420, it performs various functions of the smart phone and processes data. Optionally, the processor 1480 may include one or more processing units; preferably, the processor 1480 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1480 either.
[0175] In this embodiment, the steps executed by the processor 1480 in the terminal may be based on Figure 14 the structure shown.
[0176] The device may further include a server. Please refer to Figure 15 shown Figure 15 which is the structural diagram of the server 1500 provided in the embodiment of the present application. The server 1500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (Central Processing Units, abbreviated as CPU) 1522 (for example, one or more processors) and a memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 may be transient storage or persistent storage. The programs stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1522 may be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the server 1500.
[0177] The server 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0178] In this embodiment, the central processing unit 1522 in the server 1500 may execute the following steps:
[0179] Obtain the acquired target image, where the target image includes text information and traffic signs;
[0180] Through an information matching model, the target image is processed to obtain a matching result between the text information in the target image and a traffic sign. The information matching model is an end-to-end neural network that takes the target image as input and the matching result between the text information and the traffic sign as output.
[0181] According to one aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium is used to store program code, and the program code is used to execute the information matching method of the traffic sign described in each of the foregoing embodiments.
[0182] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the foregoing embodiments.
[0183] In the description of the present application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0184] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.
[0185] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0186] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0187] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0188] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.
Claims
1. A method for information matching of traffic signs, characterized in that, The method includes: Obtaining a collected target image, where the target image includes text information and traffic signs; Processing the target image through an information matching model to obtain a matching result between the text information and the traffic signs in the target image. The information matching model is an end-to-end neural network that takes the target image as input and the matching result between the text information and the traffic signs as output. The information matching model includes a feature extraction network, a detection network, and a relationship matching network; Among them, the feature extraction network is a first neural network that takes the target image as input and the image semantic features as output; The detection network is a second neural network that takes the output of the feature extraction network as input and the text detection result and the traffic sign detection result as output. The text detection result includes text position information indicating the position of the text information in the target image, and the traffic sign detection result includes sign position information indicating the position of the traffic sign in the target image; The relationship matching network is a third neural network that takes the output of the detection network as input and the matching result between the text information and the traffic signs in the target image as output.
2. The method according to claim 1, wherein The processing the target image through the information matching model to obtain the matching result between the text information and the traffic signs in the target image includes: Performing feature extraction on the target image through the feature extraction network to obtain image semantic features; Determining the text detection result and the traffic sign detection result through the detection network according to the image semantic features; Determining the matching result between the text information and the traffic signs in the target image through the relationship matching network according to the text position information in the text detection result and the sign position information in the traffic sign detection result.
3. The method according to claim 2, characterized in that The determining the matching result between the text information and the traffic signs in the target image through the relationship matching network according to the text position information in the text detection result and the sign position information in the traffic sign detection result includes: Determining a sign position encoding corresponding to the sign position information and a text position encoding corresponding to the text position information; Fusing the sign position encoding, the text position encoding, and the image semantic features to obtain a fused feature; Determining the matching result between the text information and the traffic signs in the target image according to the fused feature.
4. The method according to claim 3, wherein The fusing the sign position encoding, the text position encoding, and the image semantic features to obtain a fused feature includes: Determining a region of interest according to the sign position information and the text position information; Cropping the region of interest corresponding to the region of interest from the image semantic features to obtain an interested semantic feature; Fusing the interested semantic feature with the corresponding sign position encoding or text position encoding to obtain the fused feature.
5. The method according to claim 2, wherein The detection network includes a text information detection branch and a traffic sign detection branch; Among them, the text information detection branch is a fourth neural network that takes the output of the feature extraction network as input and the text detection result as output; The traffic sign detection branch is a fifth neural network that takes the output of the feature extraction network as input and the traffic sign detection result as output.
6. The method according to claim 5, wherein The determining the text detection result and the traffic sign detection result through the detection network according to the image semantic features includes: Performing feature transformation on the image semantic features through the text information detection branch to obtain first semantic features, where the significance of the features of the text information in the first semantic features is higher than the significance of the features of the text information in the image semantic features; Obtaining the text detection result through the text information detection branch according to the first semantic features; Performing feature transformation on the image semantic features through the traffic sign detection branch to obtain second semantic features, where the significance of the features of the traffic signs in the second semantic features is higher than the significance of the features of the traffic signs in the image semantic features; Obtaining the traffic sign detection result through the traffic sign detection branch according to the second semantic features.
7. The method according to any one of claims 1-6, characterized in that, The text detection result is represented by a first detection box in the target image, and the traffic sign detection result is represented by a second detection box in the target image.
8. The method according to any one of claims 2-6, characterized in that, The text detection result also includes a first confidence score corresponding to the text detection result, and the traffic sign detection result also includes a second confidence score corresponding to the traffic sign detection result. The determining the matching result between the text information and the traffic signs in the target image according to the text position information in the text detection result and the sign position information in the traffic sign detection result through the relationship matching network includes: Selecting M text detection results according to the first confidence score, and selecting N traffic sign detection results according to the second confidence score and inputting them into the relationship matching network; Outputting an affiliation relationship graph through the relationship matching network, where the affiliation relationship graph is used to represent the matching result between the text information corresponding to the M text detection results and the traffic signs corresponding to the N traffic sign detection results.
9. The method according to any one of claims 1-6, characterized in that, The method further includes: Constructing an information initial matching model, where the information initial matching model includes an initial feature extraction network, an initial detection network, and an initial relationship matching network; Obtaining training samples in the training sample set, where the training samples include training images and the true matching results between text information and traffic signs; Inputting the training images into the information initial matching model, and successively passing through the initial feature extraction network, the initial detection network, and the initial relationship matching network for processing to obtain the output content of the initial relationship matching network, where the output content includes the predicted matching result between the text information and the traffic signs; Constructing a loss function according to the predicted matching result and the true matching result; Adjust the model parameters of the initial information matching model according to the loss function, and determine the information matching model according to the adjusted model parameters when the training conditions are met and the network structure of the initial information matching model.
10. The method according to claim 9, characterized in that, The adjusting the model parameters of the initial information matching model according to the loss function, and determining the information matching model according to the adjusted model parameters when the training conditions are met and the network structure of the initial information matching model includes: Perform backpropagation of the loss represented by the loss function in the initial feature extraction network, the initial detection network, and the initial relationship matching network, and adjust the weights and model parameters of the initial feature extraction network, the initial detection network, and the initial relationship matching network according to the loss until the training conditions are met, to obtain the information matching model.
11. An information matching device for a traffic sign, characterized in that, The apparatus includes an acquisition unit and a matching unit: The acquisition unit is configured to acquire a target image obtained by collection, where the target image includes text information and a traffic sign; The matching unit is configured to process the target image through an information matching model to obtain a matching result between the text information and the traffic sign in the target image. The information matching model is an end-to-end neural network that takes the target image as an input and the matching result between the text information and the traffic sign as an output; the information matching model includes a feature extraction network, a detection network, and a relationship matching network; Wherein, the feature extraction network is a first neural network that takes the target image as an input and the image semantic feature as an output; The detection network is a second neural network that takes the output of the feature extraction network as an input and the text detection result and the traffic sign detection result as outputs. The text detection result includes text position information indicating the position of the text information in the target image, and the traffic sign detection result includes sign position information indicating the position of the traffic sign in the target image; The relationship matching network is a third neural network that takes the output of the detection network as an input and the matching result between the text information and the traffic sign in the target image as an output.
12. The device according to claim 11, wherein, The matching unit is configured to: Extract image semantic features from the target image through the feature extraction network; Determine the text detection result and the traffic sign detection result through the detection network according to the image semantic features; Determine the matching result between the text information and the traffic sign in the target image through the relationship matching network according to the text position information in the text detection result and the sign position information in the traffic sign detection result.
13. The device according to claim 12, characterized in that, The matching unit is configured to: Determine a sign position encoding corresponding to the sign position information and a text position encoding corresponding to the text position information; Fuse the sign position encoding, the text position encoding, and the image semantic feature to obtain a fused feature; Determine the matching result between the text information and the traffic sign in the target image according to the fused feature.
14. The device according to claim 13, wherein The matching unit is configured to: Determine the region of interest based on the signboard position information and the text position information; Crop the semantic feature corresponding to the region of interest from the image semantic features; Fuse the semantic feature of interest with the corresponding signboard position encoding or text position encoding to obtain the fused feature.
15. The device according to claim 12, characterized in that, The detection network includes a text information detection branch and a traffic signboard detection branch; Among them, the text information detection branch is a fourth neural network that takes the output of the feature extraction network as input and the text detection result as output; The traffic signboard detection branch is a fifth neural network that takes the output of the feature extraction network as input and the traffic signboard detection result as output.
16. The device according to claim 15, characterized in that, The matching unit is used for: Perform feature transformation on the image semantic features through the text information detection branch to obtain a first semantic feature, where the significance of the feature of the text information in the first semantic feature is higher than that of the feature of the text information in the image semantic features; Obtain the text detection result through the text information detection branch according to the first semantic feature; Perform feature transformation on the image semantic features through the traffic signboard detection branch to obtain a second semantic feature, where the significance of the feature of the traffic signboard in the second semantic feature is higher than that of the feature of the traffic signboard in the image semantic features; Obtain the traffic signboard detection result through the traffic signboard detection branch according to the second semantic feature.
17. The device according to any one of claims 11-16, characterized in that, The text detection result is represented by a first detection box in the target image, and the traffic signboard detection result is represented by a second detection box in the target image.
18. The device according to any one of claims 12 - 16, characterized in that, The text detection result also includes a first confidence score corresponding to the text detection result, and the traffic signboard detection result also includes a second confidence score corresponding to the traffic signboard detection result. The matching unit is used for: Select M text detection results according to the first confidence score, and select N traffic signboard detection results according to the second confidence score and input them into the relationship matching network; Output an affiliation graph through the relationship matching network, and the affiliation graph is used to represent the matching result between the text information corresponding to the M text detection results and the traffic signboards corresponding to the N traffic signboard detection results.
19. The device according to any one of claims 11-16, characterized in that, The device further includes: a training unit; the training unit is used for: Construct an initial information matching model, where the initial information matching model includes an initial feature extraction network, an initial detection network, and an initial relationship matching network; Obtain training samples in the training sample set, and the training samples include training images and the true matching results between text information and traffic signboards; Input the training images into the initial information matching model, and after being processed by the initial feature extraction network, the initial detection network, and the initial relationship matching network in sequence, obtain the output content of the initial relationship matching network, and the output content includes the predicted matching result between the text information and the traffic signboard; Construct a loss function based on the predicted matching result and the true matching result; Adjust the model parameters of the initial information matching model according to the loss function, and determine the information matching model according to the adjusted model parameters and the network structure of the initial information matching model when the training conditions are met.
20. The device according to claim 19, characterized in that, The training unit is configured to: Backpropagate the loss represented by the loss function in the initial feature extraction network, the initial detection network, and the initial relationship matching network, and adjust the weights and model parameters of the initial feature extraction network, the initial detection network, and the initial relationship matching network according to the loss until the training conditions are met, thereby obtaining the information matching model.
21. An information matching device for a traffic signboard, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1-10 based on the instructions in the program code.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code, and the program code is used to execute the method according to any one of claims 1-10.
23. A computer program product, characterized in that, It includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1-10.
Citation Information
Patent Citations
Traffic sign board information acquisition method and system for high-precision map production
CN110501018A