A real-time recognition method and system for mileage stake numbers

Through a lightweight 2D camera combined with object detection and text recognition algorithm, combined with dynamic programming decoding and joint inference mechanism, the accuracy and speed of the mileage station recognition algorithm in complex environments is solved, real-time and stable station information acquisition is achieved, and accurate positioning of road diseases is supported.

CN119904855BActive Publication Date: 2025-07-25SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411996088.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-07-25
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the prior art, the mileage pile identification algorithm has insufficient recognition accuracy and speed in complex environments, which cannot meet the real-time requirements of road patrols, and the signal is unstable in some areas by relying on the global positioning system.

Method used

A lightweight 2D camera is used to combine object detection and text recognition algorithms, and a mileage pile image is extracted using the object detection model, and a text recognition model extracts pile numbers. Combined with dynamic programming decoding and joint inference mechanisms, the recognition accuracy is improved through multi-frame image analysis.

Benefits of technology

In complex environments, the mileage pile number is quickly and accurately identified, reducing dependence on the global navigation system, providing stable pile number information, and improving the accuracy and efficiency of road disease positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904855B_ABST
    Figure CN119904855B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time recognition method and system for mileage markers, which relates to the technical field of road operation and maintenance, and solves the technical problem that the existing technology cannot accurately recognize the mileage marker signs in images. The present invention includes preprocessing and corresponding annotation of road foreground image data respectively to obtain a target detection data set and a text recognition data set, training a target detection algorithm model using the target detection data set to obtain a target detection model, then training a text recognition algorithm model using the text recognition data set to obtain a text recognition model, and then using the target detection model to extract mileage marker images from real-time road foreground images, and then using the text recognition model to obtain mileage markers from the mileage marker images. The present invention can quickly and accurately identify the position of mileage markers in complex road scenes, and still performs excellently even in cases such as insufficient lighting and stain occlusion, ensuring the real-time performance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road operation and maintenance, and particularly relates to a method and a system for real-time identification of mileage stake numbers. Background Art

[0002] Expressways are an important part of China's transportation infrastructure, and the automated identification technology of road surface diseases has become an important means to evaluate the service performance of expressways. However, only identifying road surface diseases is not sufficient to achieve road surface inspection. This is because road maintenance personnel also need to know the mileage stake number position where the road surface diseases are located. Currently, most road inspection vehicles are equipped with a global positioning and navigation system to record the driving route trajectory in real time. However, when using the global positioning and navigation system, the following points need to be considered: (1) The global positioning system needs to cooperate with the camera, which takes extra computing time; (2) The signal is weak on mountainous or tunnel expressways, and it is impossible to ensure the real-time update of longitude and latitude information; (3) The longitude and latitude coordinates provided by the global positioning and navigation system are not sufficient to accurately and intuitively provide the mileage position information of road surface diseases. Mileage stake plates are laid on both sides of expressways and are natural positioning references. Even in mountainous areas or tunnels, mileage stake plates can provide intuitive and accurate positioning references for road maintenance personnel. In addition, mileage stake plates can also be used as an auxiliary means for the global positioning system to obtain more reliable and accurate mileage positions of road surface diseases. Therefore, the accurate identification of mileage stake number positions and the accurate extraction of mileage stake number texts are crucial for the positioning of road surface diseases.

[0003] However, within the scope of the literature investigated by the research team, during the acquisition process by road inspection vehicles, most adopt the method of manually recording the mileage stake number positions. This manual recording method is time-consuming and laborious, and the recorder needs to maintain concentration at all times. The existing deep learning-based mileage stake number recognition algorithms generally have disadvantages such as low recognition accuracy, slow recognition speed, and inability to handle various emergencies that occur in the actual environment. Therefore, there is an urgent need for an automated mileage stake number extraction technology to replace the time-consuming and laborious manual operation and be able to handle various situations that occur in the actual application process (such as missing stakes, unclear stake numbers) to improve the recognition accuracy and applicability. Based on this, the present invention aims to provide an automated and intelligent real-time mileage stake number information extraction method to provide reliable mileage information for the positioning of road surface diseases.

[0004] Using the road foreground image captured by a lightweight 2D camera, the present invention automatically extracts the position and text information of mileage markers from the image by using object detection algorithms and text recognition algorithms. Since mileage marker plates often appear as small targets in the image, it is required that the object detection algorithm can accurately identify small target objects from the complex road traffic background. In addition, in the case of insufficient lighting, smear occlusion, or long distance, the mileage marker may not be clear enough. Therefore, it is also necessary to combine multiple frames of images for joint analysis and inference to improve the reliability of recognition. The sizes and text settings of mileage markers on different highways are also different, specifically reflected in the font, font size, and layout angle of the mileage markers. Therefore, the text recognition algorithm also needs to have strong generalization ability to be able to process mileage markers in different states and accurately extract the mileage marker text sequence.

[0005] Existing object detection algorithms have poor robustness in detecting small targets, resulting in the inability to accurately identify mileage marker plates in images. Summary of the Invention

[0006] To solve the problems existing in the above-mentioned prior art, the present invention provides a real-time recognition method and system for mileage markers, which solves the technical problem that the prior art cannot accurately identify mileage marker plates in images.

[0007] A real-time recognition method for mileage markers includes: preprocessing and corresponding annotation of road foreground image data respectively to obtain an object detection data set and a text recognition data set, training an object detection algorithm model using the object detection data set to obtain an object detection model, and then training a text recognition algorithm model using the text recognition data set to obtain a text recognition model. Then, using the object detection model to extract mileage marker images from real-time road foreground images, and then using the text recognition model to obtain mileage markers from the mileage marker images; the object detection algorithm model adopts two parallel branches, one branch extracts image features, and the other branch encodes the image features in terms of position encoding.

[0008] Further, the extraction of image features includes: preprocessing road foreground images with the same size, then using a residual network to extract features, and then performing dynamic position encoding on the extracted feature layers. On the basis of the original position encoding, a learnable position weight coefficient is added to each position, so that it can automatically learn the relative position relationship during backpropagation.

[0009] Further, the position encoding of the image features includes: using an enhanced feature encoding network with a gating mechanism to further capture effective features, and then using a feature decoding network to decode the enhanced feature layer obtained from the enhanced feature encoding network to obtain useful mileage stake number feature position information and category information. Finally, using a regression prediction head and a classification prediction head to further output the feature layer obtained in the feature decoding stage as available coordinate and category information.

[0010] Further, the step of using an enhanced feature encoding network with a gating mechanism to further capture effective features includes:

[0011] For the feature layer with position encoding obtained in the feature extraction stage, first pass it into the proposed gating mechanism, whose main function is to calculate a gating coefficient for each feature channel as a neuron parameter. When a certain channel has a positive effect on the final stake number positioning, it will pass through the set gating mechanism; otherwise, it will be filtered out.

[0012] The feature layer passing through the gating mechanism will be further passed into the multi-head self-attention mechanism to capture the long-range dependence relationship between feature information.

[0013] After the Add operation and layer normalization, it further enters the fully connected layer to obtain the retrieved effective feature layer. Finally, through the Add and normalization operations, it is passed into the feature decoding network to further realize the recovery of the feature layer.

[0014] Further, the training to obtain the object detection model includes: training the object detection algorithm model based on the Hungarian algorithm. The work of the Hungarian algorithm is to match multiple prediction results with multiple ground truth boxes. This process requires calculating a cost matrix, and then according to the cost matrix, using the Hungarian algorithm to calculate the situation with the lowest cost. Among them, the calculation process of the cost matrix is as follows:

[0015] (a) Calculate the classification cost; (b) Calculate the L1 cost between the predicted box and the ground truth box; (c) Calculate the IOU cost between the predicted box and the ground truth box; Add the three calculation results according to the weights to obtain the cost matrix.

[0016] Further, the text recognition algorithm model includes: using EfficientNet to extract features to generate a stake number feature sequence and input it into a GRU recurrent neural network for mileage stake number prediction.

[0017] Furthermore, the EfficientNet includes: EfficientNet-B0 as the baseline network. The first stage is a 3×3 convolutional layer with a stride of 2 for feature extraction. The second to seventh stages use inverted residual blocks (MBConv). The MBConv block contains a 1×1 convolution (for expanding the number of channels), a 3×3 depthwise separable convolution, and another 1×1 convolution to compress the number of channels. These blocks also include residual connections and Squeeze-and-Excitation (SE) modules. The last stage contains a 1×1 convolutional layer for further feature fusion.

[0018] Furthermore, after the text recognition model obtains the mileage stake number from the mileage stake image, it uses a joint inference mechanism to confirm the content of the ambiguous stake number based on the front and rear stake number information. The joint inference mechanism includes: comparing and confirming by combining the recognition information of the front and rear frames, using a confidence threshold judgment to correct the uncertain stake number prediction value. If the recognition confidence of a certain stake number is lower than the threshold, it is reconfirmed by referring to the front and rear frame data to ensure that the final output result is more accurate. Furthermore, the joint inference mechanism also includes:

[0019] Multi-frame image fusion and joint inference: By using multiple frames of images continuously captured by the detection vehicle and fusing the multi-frame information, the current stake number can be inferred more accurately;

[0020] Up and down line logic verification and dynamic programming decoding: By constructing a state graph and a transition cost function to represent the logical relationship between stake numbers, each state represents a possible stake number prediction value, the nodes in the state graph represent the possibility of stake numbers, the edges represent the logical transition from one stake number to another, and the transition cost function is used to judge the rationality of the transition from one stake number to the next. The dynamic programming algorithm is used to find the path with the minimum cost in the state graph to ensure that the stake number sequence conforms to logical continuity and real distribution;

[0021] Time series consistency analysis: By smoothing the prediction results of consecutive multi-frame stake numbers, the fluctuations in single-frame image recognition are eliminated.

[0022] A real-time recognition system for mileage stake numbers adopts a real-time recognition method for mileage stake numbers, including a data acquisition and processing module, a target detection module, a text recognition module, and an output module. The data acquisition module is used for real-time acquisition and preprocessing of detection target images. The target detection module includes a target detection model for determining the mileage stake image based on the preprocessed detection target image. The text recognition module includes a text recognition model for determining the mileage stake number based on the mileage stake image.

[0023] The beneficial effects of the present invention include:

[0024] In terms of algorithm innovation, the present invention proposes an object detection algorithm and a text recognition algorithm, and combines dynamic programming decoding with a joint inference mechanism, significantly improving the detection ability for small targets and blurred stake numbers and the adaptability to handle complex scenarios. The object detection algorithm optimizes the feature extraction and detection process, and can quickly and accurately identify the position of the mileage stake number in complex road scenarios. Even in the case of insufficient lighting, stain occlusion, etc., it still performs excellently, ensuring the real-time performance of the system. The text recognition algorithm does not require labeling each character one by one, greatly reducing the time and cost of data preparation, and ensuring that the recognized stake number sequence conforms to the actual logical distribution through the dynamic programming decoding mechanism, further enhancing the robustness and adaptability of the algorithm under different lighting and complex backgrounds.

[0025] Another major innovation of the present invention is the introduction of a joint inference mechanism. By combining the information of multiple frames of images, it jointly infers and confirms blurred stake numbers. In practical applications, the stake number often becomes unclear due to lighting changes, stain occlusion, or long distance, and a single-frame image may not be able to accurately identify it. The joint inference mechanism analyzes the stake number information of the front and rear frames, and uses time series consistency and logical verification to significantly improve the recognition success rate and accuracy, enabling the system to still provide high-quality stake number information in real time and stably in complex environments.

[0026] At the level of engineering practical application, the mileage stake number automatic extraction algorithm of the present invention realizes real-time and stable acquisition of mileage stake number information during the inspection process of road inspection vehicles, avoiding the dependence on the Global Navigation Satellite System (GNSS). Especially in areas where GNSS signals are unstable, such as tunnels or mountainous areas, it can still provide accurate stake number position information. This provides reliable support for the precise positioning of road diseases, enabling maintenance personnel to quickly and intuitively determine the disease location and take efficient repair measures. In addition, the lightweight 2D camera solution of the present invention combined with an efficient algorithm design provides a lightweight and low-cost automatic extraction method, which is convenient to be deployed on ordinary inspection vehicles, thus expanding the application scope of the technology and improving the overall efficiency and quality of road maintenance. Brief Description of the Drawings

[0027] Figure 1 It is a flowchart of the integrated mileage stake number real-time recognition method involved in the embodiment of the present application.

[0028] Figure 2 It is a schematic diagram of the object detection algorithm model involved in the embodiment of the present application.

[0029] Figure 3 It is a schematic diagram of the multi-head attention mechanism involved in the embodiment of the present application.

[0030] Figure 4 It is a schematic diagram of the gating mechanism involved in the embodiment of the present application.

[0031] Figure 5 This is a schematic diagram of the text recognition algorithm model involved in the embodiments of the present application.

[0032] Figure 6 This is a real-time road foreground image involved in the embodiments of the present application.

[0033] Figure 7 This is the detection result of the real-time road foreground image involved in the embodiments of the present application. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0035] A real-time recognition method for mileage stake numbers, as Figure 1 shown, includes: preprocessing and corresponding annotation of road foreground image data respectively to obtain an object detection data set and a text recognition data set, training an object detection algorithm model using the object detection data set to obtain an object detection model, then training a text recognition algorithm model using the text recognition data set to obtain a text recognition model, and then using the object detection model to extract mileage stake images from the real-time road foreground image, and then using the text recognition model to obtain mileage stake numbers from the mileage stake images; the object detection algorithm model adopts two parallel branches, one branch extracts image features, and the other branch performs position encoding on the image features in a position encoding manner.

[0036] Specifically, the process of obtaining road foreground image data is as follows:

[0037] The road foreground image is obtained by using the front industrial 2D camera installed on the road inspection vehicle. The camera has high resolution and high dynamic range, and can capture clear images under various lighting and weather conditions. In order to ensure the accuracy and reliability of the data, the camera is precisely calibrated to ensure the quality and accuracy of the image. The collected road foreground image data is used to build a mileage pile number database. This database not only contains the images of mileage pile numbers, but also records the corresponding geographic location information and timestamps, providing important basic information for subsequent data processing and analysis. Since different highways often have different pile numbers, it is necessary to obtain the pile number data of multiple different highways to ensure data diversity in order to build a sample database for training algorithm models. The specific roads for data collection include but are not limited to various national highway kilometers, provincial highways, ring highways, urban expressways, and other sections with pile numbers, all of which are within the scope of data collection.

[0038] In the data annotation stage, a semi-automatic data annotation method is used:

[0039] The location and related information of the mileage posts are manually identified and marked in the image for preliminary annotation, and then the initial deep learning model is trained using the annotated data. Once the deep learning model reaches a certain accuracy, the data is automatically annotated through the deep learning model to improve the annotation efficiency. At the same time, in order to ensure the quality of the final data, all the data annotated by the model will be manually reviewed. This process ensures the accuracy and consistency of the data set while ensuring the efficiency of database establishment, providing high-quality input data for the training of the deep learning model.

[0040] The specific implementation steps are as follows:

[0041] (1) Obtain road foreground images of national highways, provincial highways, ring roads, urban expressways, and other road sections with stake numbers, and build an original image data platform;

[0042] (2) In view of the diversity of mileage post signs on different highways, it is necessary to select typical image samples with mileage post signs to be identified from the constructed raw data platform and try to include all the image samples with mileage post signs to be identified;

[0043] (3) Based on the selected typical samples, manually label a batch of samples. Based on this batch of labeled samples, first train a deep learning target detection model with preliminary recognition capabilities;

[0044] (4) Using this model with preliminary recognition capabilities, the remaining data can be automatically labeled and then manually calibrated. The efficiency of such semi-automatic labeling will be much higher than that of full manual labeling.

[0045] In another embodiment, the extraction of image features includes: preprocessing road foreground images with consistent dimensions, then using a residual network to extract features, and then performing dynamic position encoding on the extracted feature layer. On the basis of the original position encoding, a learnable position weight coefficient is added to each position, enabling it to automatically learn the relative position relationship during backpropagation.

[0046] Specifically, it includes: preprocessing of target detection images

[0047] In the stage of mileage stake number target detection, image preprocessing is a key step. First, denoise, adjust the contrast, and balance the color of the captured road image to improve the visibility of the mileage stake number in the image. Then, use Gaussian blur and edge enhancement techniques to further highlight the features of the mileage stake number, which is particularly important for subsequent target detection. In addition, considering the influence of different lighting conditions and weather changes, adaptive histogram equalization processing is introduced to ensure that the image can maintain consistent quality in various environments. Finally, scale the road image data with the stake number to obtain a fixed pixel size of 640×640 and input it into the algorithm model to improve the processing speed and meet the real-time requirement.

[0048] Propose an efficient target detection algorithm model to identify and locate the mileage stake number, and its model structure is as Figure 2 shown. Its model structure is specially designed to enhance the detection ability of small targets, especially for cases such as blurred stake numbers, insufficient lighting, and occlusion.

[0049] Feature extraction and dynamic position encoding

[0050] In order to improve the detection ability of small targets and blurred stake numbers, especially for insufficient lighting or occlusion in complex road scenes, this part mainly includes two parallel branches. The first branch uses a ResNet50 residual network structure to extract image features, and the second branch uses the Transformer position encoding method to perform position encoding on the image features. Because in the Transformer algorithm adopted in the present invention, not only feature extraction but also feature query is required. Therefore, position information (Positioning Embedding) needs to be added to all features so that the network has the ability to distinguish different regions. The specific implementation steps are as follows:

[0051] (1) Input color image data with a size of 640×640×3 (width × height × channels) and perform image preprocessing, specifically the normalization operation of the image data, mapping all pixel points between 0 and 1;

[0052] (2) Use the ResNet50 feature extraction backbone network to perform preliminary feature extraction, providing a refined feature layer for the subsequent self-attention encoding;

[0053] (3) Different from existing deep learning-based object detection algorithms, the present invention proposes to perform dynamic position encoding on the feature layer extracted by ResNet50, enabling it to have the ability to distinguish different feature regions and not lose corresponding position information during subsequent object query and object recovery processes. The specific details are that on the basis of the original position encoding, a learnable position weight coefficient is added to each position, enabling it to automatically learn the relative position relationship during backpropagation, which helps to enhance the attention to the positions of small target objects.

[0054] In another embodiment, the further capturing of effective features by using the enhanced feature encoding network with a gating mechanism includes:

[0055] For the feature layer with position encoding obtained in the feature extraction stage, first pass it into the proposed gating mechanism, as Figure 4 shown. Its main function is to calculate a gating coefficient for each feature channel as a neuron parameter; when a certain channel has a positive effect on the final stake number positioning, it will pass through the set gating mechanism, otherwise it will be filtered out;

[0056] The feature layer passing through the gating mechanism will be further passed into the multi-head self-attention mechanism, as Figure 3 shown, for capturing the long-range dependence relationship between feature information;

[0057] After the Add operation and layer normalization, it further enters the fully connected layer to obtain the effective feature layer after retrieval. Finally, through the Add and Normalization operations, it is passed into the feature decoding network to further realize the recovery of the feature layer.

[0058] Use the feature decoding network to perform decoding operations on the enhanced feature layer obtained in the enhanced feature encoding network to obtain useful stake number feature position information and category information, realizing end-to-end object detection.

[0059] (1) Object query: Use a learnable query vector q to query the enhanced effective feature layer to obtain the prediction result;

[0060] (2) After passing through the multi-head self-attention mechanism structure, use the refined feature layer as Q, and use the feature matrix extracted by the enhanced feature encoding network as K and V, and then pass them into the multi-head self-attention mechanism for another feature enhancement and feature retrieval.

[0061] The result prediction head (which includes a regression prediction head and a classification prediction head) further outputs the feature layer obtained in the feature decoding stage into available coordinate and category information.

[0062] (1) Regression prediction head: It includes fully connected layers, ReLU activation functions, fully connected layers, ReLU activation functions, fully connected layers, and Sigmoid activation functions. The final output contains four coefficients, namely the coordinates of the center point of the predicted bounding box and the width and height of the predicted bounding box. Among them, 100 predicted bounding boxes are used.

[0063] (2) Classification prediction head: It uses a single fully connected layer, and the final output contains two coefficients. Since only one target is recognized, the output is the number of classes + 1.

[0064] In another embodiment, the training to obtain the object detection model includes: training the object detection algorithm model based on the Hungarian algorithm. The work of the Hungarian algorithm is to match multiple prediction results and multiple ground truth boxes. This process requires calculating a cost matrix, and then according to the cost matrix, using the Hungarian algorithm to calculate the case of the lowest cost. Among them, the calculation process of the cost matrix is as follows:

[0065] (a) Calculate the classification cost; (b) Calculate the L1 cost between the predicted bounding box and the ground truth box; (c) Calculate the IOU cost between the predicted bounding box and the ground truth box; Add the three calculation results according to the weights to obtain the cost matrix.

[0066] Specifically, in the object detection algorithm, during training, the matching process of positive samples is based on the Hungarian algorithm. The work of the matching algorithm is to match 100 prediction results and N ground truth boxes. One ground truth box only matches one prediction result, and the other prediction results are used as the background for fitting. Therefore, the work of the matching algorithm is to find the N prediction results that are most suitable for predicting N ground truth boxes. Therefore, it is necessary to calculate a cost matrix (Cost matrix) to represent the relationship between 100 prediction results and N ground truth boxes. This cost matrix consists of three parts:

[0067] (a) Calculate the classification cost. Among the obtained prediction results, obtain the predicted value corresponding to the class of this ground truth box. If the predicted value is larger, it means that this predicted bounding box is predicted more accurately, and its cost is lower; (b) Calculate the L1 cost between the predicted bounding box and the ground truth box. Among the obtained prediction results, obtain the coordinates of the predicted bounding box, and calculate an L1 distance between the coordinates of the predicted bounding box and the coordinates of the ground truth box. The more accurate the prediction, the lower its cost; (c) Calculate the IOU cost between the predicted bounding box and the ground truth box. Among the obtained prediction results, obtain the coordinates of the predicted bounding box, and calculate an IOU distance between the coordinates of the predicted bounding box and the coordinates of the ground truth box. The more accurate the prediction, the lower its cost. The above three parts are added according to certain weights to obtain the cost matrix, and then according to the cost matrix, use the Hungarian algorithm to calculate the case of the lowest cost.

[0068] In this implementation, for the proposed object detection model, tests were conducted on 2,000 actual measured road images with mileage stake numbers, compared with other current advanced object detection algorithms, namely YOLOv4-m, YOLOv5-m, FCOS, DETR, Faster R-CNN, and SSD. The test environment was Python 3.8, the deep learning platform was TensorFlow 2.0, and the GPU hardware configuration was NVIDIA RTX 3070. The test metric results are shown in Table 1 below.

[0069] Table 1 Performance of Different Object Detection Algorithms

[0070]

[0071] Among them, the evaluation metrics adopted the three most representative metrics in the current field of object detection intelligent algorithms, namely F1-Score and mAP. The larger the values of these two metrics, the better the performance of the algorithm model. The metrics of all algorithms were measured at a confidence level of 0.5. Compared with the current mainstream object detection algorithms: YOLOv4-m, YOLOv5-m, FCOS, DETR, Faster R-CNN, and SSD, the algorithm proposed in the present invention has obvious advantages in the recognition of small objects of mileage stake numbers.

[0072] In another embodiment, the text recognition algorithm model is an efficient text recognition algorithm for extracting text information from the detected mileage stake number logo images. Its detailed architecture is as Figure 5 shown. It includes: using EfficientNet to extract features to generate a stake number feature sequence and inputting it into a GRU recurrent neural network for mileage stake number prediction.

[0073] In the mileage stake number text recognition stage, the cropped stake number images are unified to a size of 128×64 and fed into the text recognition model.

[0074] GRU is a recurrent neural network (RNN) for processing sequential data, aiming to solve the problem of gradient disappearance faced by traditional RNNs when processing long sequential data. Through special structural design, GRU can more effectively capture long-term dependencies in sequential modeling.

[0075] (1) Update gate: Determines to what extent information needs to be passed from the past to the future. This helps the model decide how much past information to retain in the current state.

[0076] (2) Reset gate: Decides how much past information needs to be forgotten. This helps the model decide how much previous information to discard when calculating the candidate for the current state.

[0077] (3) Candidate hidden state: Use a reset gate to determine how much past information should be forgotten. Combine the current input and the past hidden state (adjusted by the reset gate) to generate the current candidate hidden state.

[0078] The present invention uses GRU for the prediction of the mileage post sequence, mainly taking advantage of its effective handling of the vanishing gradient problem. In addition, the GRU structure is simpler, has fewer parameters, so it is faster to train and more effective and fast in practical applications.

[0079] In another embodiment, the EfficientNet includes: EfficientNet-B0 as the baseline network. The first stage is a 3×3 convolutional layer with a stride of 2 for feature extraction; the second to seventh stages use inverted residual blocks (MBConv). The MBConv block includes a 1×1 convolution (for expanding the number of channels), a 3×3 depthwise separable convolution, and another 1×1 convolution to compress the number of channels. These blocks also include residual connections and Squeeze-and-Excitation (SE) modules. The last stage includes a 1×1 convolutional layer for further feature fusion.

[0080] In another embodiment, after the text recognition model obtains the mileage post number from the mileage post image, it uses a joint inference mechanism to confirm the content of the ambiguous mileage post number based on the information of the front and rear mileage post numbers. The joint inference mechanism includes: comparing and confirming by combining the recognition information of the front and rear frames, and using a confidence threshold judgment to correct the uncertain mileage post number prediction value. If the recognition confidence of a certain mileage post number is lower than the threshold, it is reconfirmed by referring to the front and rear frame data to ensure that the final output result is more accurate.

[0081] The mileage post joint inference mechanism aims to improve the recognition accuracy and robustness of the system for mileage post numbers in complex environments, especially in the face of ambiguous mileage post numbers, insufficient lighting, smear occlusion, etc., to ensure the rationality and stability of the recognition results. This mechanism combines methods such as dynamic programming decoding, multi-frame image joint inference, and time series consistency analysis to achieve the best decoding and inference of the mileage post sequence.

[0082] (1) Multi-frame image fusion and joint inference: To improve the recognition ability for ambiguous mileage post numbers, this mechanism uses multiple frames of images continuously captured by the detection vehicle. These multiple frames of images provide mileage post number information at different time points. Especially when the mileage post number is partially occluded or the lighting is uneven, by fusing the multi-frame information, the current mileage post number can be inferred more accurately. This method effectively improves the detection success rate for ambiguous mileage post numbers by combining the recognition results of multiple mileage post numbers in the spatio-temporal dimension.

[0083] (2) Up - and - down logical verification and dynamic programming decoding: The mileage markers on the highway are arranged in a fixed logical order, usually increasing or decreasing. In the dynamic programming decoding of mileage marker features, the logical relationship between mileage markers is represented by constructing a state diagram and a transition cost function.

[0084] State diagram construction: Each state represents a possible predicted value of the mileage marker. The nodes in the state diagram represent the possibilities of mileage markers, and the edges represent the logical transitions from one mileage marker to another.

[0085] Define transition cost: The transition cost function is used to judge the rationality of the transition from one mileage marker to the next. For example, mileage markers should be arranged in a fixed increasing or decreasing order. A larger transition cost indicates an illogical jump or incorrect order.

[0086] Dynamic programming to find the optimal path: Use the dynamic programming algorithm to find the path with the minimum cost in the state diagram, ensuring the logical continuity and realistic distribution of the mileage marker sequence. This decoding mechanism effectively filters out unreasonable mileage marker jumps or outliers and outputs the optimal decoded sequence that conforms to the logic.

[0087] (3) Time - series consistency analysis: Time - series consistency analysis is also introduced in the joint inference mechanism. By smoothing the predicted results of mileage markers in multiple consecutive frames, the fluctuations in single - frame image recognition are eliminated. Specifically, through the time - weighted average method, the mileage marker recognition results of consecutive frames are fused to obtain a more stable and consistent mileage marker sequence. This smoothing of the time series can reduce the recognition errors caused by single - frame image quality problems.

[0088] (4) Secondary confirmation of fuzzy mileage markers: In some cases, the recognition results of single - frame mileage markers may be uncertain. For this reason, the joint inference mechanism introduces the secondary confirmation of fuzzy mileage markers. By comparing and confirming the recognition information of the front and back frames and using a confidence threshold judgment to correct the uncertain predicted values of mileage markers. If the recognition confidence of a certain mileage marker is lower than the threshold, it is re - confirmed by referring to the data of the front and back frames to ensure that the final output result is more accurate.

[0089] Through steps such as multi - frame image fusion, up - and - down logical verification, dynamic programming decoding, time - series consistency, and secondary confirmation, the finally output mileage marker recognition results not only conform to the arrangement rules of highway mileage markers logically but also ensure the stability and accuracy of the prediction in both spatial and temporal dimensions. This comprehensive mechanism significantly improves the recognition accuracy of the system in complex environments, especially in cases where mileage markers are fuzzy or information is missing, and can effectively perform joint inference and correction.

[0090] In this embodiment, the text recognition algorithm proposed by the present invention is used to test on 2,000 test images together with the currently advanced text recognition algorithms EasyOCR and PaddleOCR. The performance indicators of the test are shown in Table 2 below.

[0091] Table 2 Performance of Different Text Recognition Algorithms

[0092]

[0093]

[0094] Among them, the accuracy evaluation index is the accuracy rate that the mileage stake number sequence of the currently predicted stake number image is all predicted correctly. Compared with the currently mainstream EasyOCR and PaddleOCR, the text recognition algorithm proposed by the present invention is more robust in the extraction of the mileage stake number sequence.

[0095] Specifically, the detection result of detecting the real-time foreground road by using a real-time recognition method for mileage stake numbers involved in this embodiment is as Figures 6-7 shown.

[0096] A real-time recognition system for mileage stake numbers includes a data acquisition and processing module, a target detection module, a text recognition module, and an output module. The data acquisition module is used for real-time acquisition and preprocessing of the detection target image. The target detection module includes a target detection model for determining the mileage stake image based on the preprocessed detection target image. The text recognition module includes a text recognition model for determining the mileage stake number based on the mileage stake image.

[0097] The above embodiments only represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.

Claims

1. A real-time recognition method for mileage stake numbers, characterized in that, Including: Preprocessing and corresponding annotation are respectively performed on road foreground image data to obtain an object detection data set and a text recognition data set. The object detection algorithm model is trained using the object detection data set to obtain an object detection model, and then the text recognition algorithm model is trained using the text recognition data set to obtain a text recognition model. Then, the object detection model is used to extract mileage stake images from real-time road foreground images, and the text recognition model is used to obtain mileage stake numbers from the mileage stake images; the object detection algorithm model adopts two parallel branches, one branch extracts image features, and the other branch performs position encoding on the image features in a position encoding manner; After the text recognition model obtains the mileage stake number from the mileage stake image, a joint inference mechanism is used to confirm the content of the fuzzy stake number based on the front and rear stake number information. The joint inference mechanism includes: comparing and confirming by combining the recognition information of the front and rear frames, using a confidence threshold judgment to correct the uncertain stake number prediction value. If the recognition confidence of a certain stake number is lower than the threshold, it is reconfirmed by referring to the front and rear frame data to ensure that the finally output result is more accurate; The joint inference mechanism further includes: Multi-frame image fusion and joint inference: Using multiple frames of images continuously captured by a detection vehicle, the current stake number is inferred by fusing multi-frame information; Up and down link logical verification and dynamic programming decoding: The logical relationship between stake numbers is represented by constructing a state graph and a transition cost function. Each state represents a possible stake number prediction value. The nodes in the state graph represent the possibility of stake numbers, and the edges represent the logical transition from one stake number to another. The transition cost function is used to judge the rationality of the transition from one stake number to the next. The dynamic programming algorithm is used to find the path with the minimum cost in the state graph to ensure that the stake number sequence conforms to logical continuity and real distribution; Time series consistency analysis: By smoothing the prediction results of consecutive multi-frame stake numbers, the fluctuations of single-frame image recognition are eliminated.

2. The real-time identification method of mileage stake number according to claim 1, wherein The extraction of the image features includes: preprocessing road foreground images with consistent sizes, then using a residual network to extract features, and then performing dynamic position encoding on the extracted feature layers. On the basis of the original position encoding, a learnable position weight coefficient is added to each position so that it can automatically learn the relative position relationship during backpropagation.

3. The real-time recognition method of mileage stake number according to claim 1, characterized in that, The position encoding of the image features includes: using an enhanced feature encoding network with a gating mechanism to further capture effective features, and then using a feature decoding network to perform decoding operations on the enhanced feature layer obtained from the enhanced feature encoding network to obtain useful mileage stake number feature position information and category information. Finally, the regression prediction head and the classification prediction head are used to further output the feature layer obtained in the feature decoding stage as available coordinate and category information.

4. The real-time recognition method of mileage stake number according to claim 3, characterized in that The use of an enhanced feature encoding network with a gating mechanism to further capture effective features includes: First, the extracted image features are fed into the proposed gating mechanism to calculate a gating coefficient for each feature channel as neuron parameters. When a certain channel has a positive effect on the final station number positioning, it will pass through the set gating mechanism; otherwise, it will be filtered out. The feature layer passing through the gating mechanism will be further fed into the multi-head self-attention mechanism to capture the long-range dependence relationships between feature information. After the Add operation and layer normalization Normalization, it further enters the fully connected layer to obtain the effective feature layer after retrieval. Finally, through the Add and Normalization operations, it is fed into the feature decoding network to further restore the feature layer.

5. The real-time identification method of mileage stake number according to claim 1, characterized in that, The training to obtain the object detection model includes: training the object detection algorithm model based on the Hungarian algorithm. The work of the Hungarian algorithm is to match multiple prediction results and multiple ground truth boxes. This process requires calculating the cost matrix, and then according to the cost matrix, using the Hungarian algorithm to calculate the situation with the lowest cost. Among them, the calculation process of the cost matrix is as follows: (a) Calculate the classification cost; (b) Calculate the L1 cost between the predicted box and the ground truth box; (c) Calculate the IOU cost between the predicted box and the ground truth box; Add the three calculation results according to the weights to obtain the cost matrix.

6. The real-time recognition method of mileage stake number according to claim 1, characterized in that, The text recognition algorithm model includes: using EfficientNet to extract features to generate a station number feature sequence and input it into the GRU recurrent neural network for predicting the mileage station number.

7. A real-time recognition method for mileage stake numbers according to claim 6, characterized in that, The EfficientNet includes: EfficientNet-B0 as the baseline network. The first stage is a 3×3 convolutional layer with a stride of 2 for feature extraction. The second to seventh stages use the inverted residual block MBConv. The MBConv block contains a 1×1 convolution, a 3×3 depthwise separable convolution, and another 1×1 convolution to compress the number of channels. It also includes a residual connection and an SE module, and finally includes a 1×1 convolutional layer for further fusing features.

8. A real-time identification system for mileage stake numbers, characterized in that, Adopt a real-time recognition method for mileage station numbers according to any one of claims 1-7, including a data acquisition and processing module, an object detection module, a text recognition module, and an output module. The data acquisition module is used for real-time acquisition and preprocessing of the detection target image. The object detection module includes an object detection model for determining the mileage stake image based on the preprocessed detection target image. The text recognition module includes a text recognition model for determining the mileage station number based on the mileage stake image.

Citation Information

Patent Citations

  • Road inspection target road mileage stake mark positioning method

    CN115798206A

  • Highway stake mark matching method based on binocular stereoscopic vision

    CN116721408A