Line facility ledger data processing method and device
By optimizing the line facility image object detection model of the YOLOv5 network structure, combined with the SENet network and EIoU loss function, the problem of inaccurate line facility ledger data is solved, more accurate target detection and data extraction is achieved, and real-time management and data update of line facilities are supported.
Patent Information
- Application Number
- CN202510077183.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology cannot accurately locate the target detection area, resulting in inaccurate ledger data of line facilities and making it difficult to achieve rapid positioning and precise management of facilities.
The line facility image object detection model based on the optimized YOLOv5 network structure is adopted, and the object detection and ledger data extraction are carried out in combination with the SENet network and the EIoU loss function. Images are acquired through linear array cameras, text information is extracted using OCR technology, ledger data is determined, and stored in the database.
It improves the positioning accuracy of the target detection area of the line facilities, enhances the accuracy of the ledger data, and provides reliable data support for the real-time management of line facilities and mileage correction.
Smart Images

Figure CN120032103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of line detection technology, and in particular to a method and device for processing ledger data of line facilities. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No admission is made that the description herein is prior art by inclusion in this section.
[0003] With the rapid development of the railway transportation industry, the number and types of line facilities are increasing. Various facilities including turnouts, signal machines, contact network equipment, power supply facilities, etc. are crucial to ensure the safe operation of railways. However, the traditional line facility management method mainly relies on manual ledger records or non-digital methods, which has problems such as untimely information updates, incomplete data, and low management efficiency. In addition, due to the wide coverage and complex environment of railway lines, it is difficult to achieve rapid positioning and precise management of facilities, which brings great challenges to inspection and maintenance work. In recent years, with the continuous development of machine vision technology, new solutions have been provided for the digital management of line facilities. At present, the existing technology cannot accurately locate the target detection area, resulting in inaccurate ledger data of line facilities. Therefore, there is an urgent need for a method for processing the ledger data of line facilities to solve the above problems. Summary of the invention
[0004] An embodiment of the present invention provides a method for processing ledger data of line facilities, which is used to improve the accuracy of ledger data and provide data support for real-time management and mileage correction of line facilities. The method includes:
[0005] In the target line, obtain line facility images collected by the line array camera;
[0006] Inputting the line facility image into the line facility image target detection model, and outputting the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5;
[0007] In the target detection area, OCR technology is used to extract text information from line facility images, and the record data of line facilities is determined based on the text information;
[0008] The ledger data is uniformly stored in a pre-built database.
[0009] The embodiment of the present invention further provides a data processing device for the ledger of line facilities, which is used to improve the accuracy of the ledger data and provide data support for the real-time management and mileage correction of line facilities. The device includes:
[0010] A line facility image acquisition module is used to acquire line facility images captured by a linear array camera in a target line;
[0011] The target detection area determination module is used to input the line facility image into the line facility image target detection model and output the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5;
[0012] The ledger data determination module is used to extract text information from line facility images using OCR technology within the target detection area, and determine the ledger data of the line facilities based on the text information;
[0013] The ledger data storage module is used to uniformly store ledger data in a pre-built database.
[0014] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for processing the ledger data of the line facilities when executing the computer program.
[0015] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned line facility ledger data processing method.
[0016] An embodiment of the present invention also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned line facility ledger data processing method.
[0017] In the embodiment of the present invention, the line facility image captured by the linear array camera is obtained in the target line; the line facility image is input into the line facility image target detection model, and the target detection area of the line facility is output; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the original loss function of YOLOv5 is optimized by the EIoU loss function; in the target detection area, the text information of the line facility image is extracted by the OCR technology, and the account data of the line facility is determined according to the text information; the account data is uniformly stored in a pre-built database. In the above process, the embodiment of the present invention is based on the optimized target detection YOLOv5 model to obtain the target detection area of the line facility, improve the accuracy of the positioning of the target detection area, and use the OCR technology to determine the text information of the line facility image in the target detection area as the account data, thereby improving the accuracy of the account data, and providing data support for the real-time management and mileage correction of the line facilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0019] Figure 1 It is a flow chart of a method for processing account data of line facilities in an embodiment of the present invention;
[0020] Figure 2 A flow chart of establishing a line facility image target detection model in an embodiment of the present invention;
[0021] Figure 3 This is a flow chart of obtaining ledger data through line facility identification in an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of a ledger data processing device for line facilities in an embodiment of the present invention;
[0023] Figure 5 Schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0025] Figure 1 Flow chart of a method for processing account data of line facilities in an embodiment of the present invention, the method comprising:
[0026] Step 101, in a target line, obtaining a line facility image captured by a line array camera;
[0027] Step 102, inputting the line facility image into the line facility image target detection model, and outputting the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5;
[0028] Step 103, within the target detection area, extract text information from the line facility image using OCR technology, and determine the line facility ledger data based on the text information;
[0029] Step 104, the ledger data is uniformly stored in a pre-built database.
[0030] Each step is described in detail below.
[0031] In step 101, in a target line, a line facility image captured by a line array camera is obtained.
[0032] In a specific embodiment, a linear array camera is installed at the bottom of the train and combined with a high frame rate image acquisition device to capture images of line facilities along the line during the operation of the train.
[0033] In step 102, the line facility image is input into the line facility image target detection model, and the target detection area of the line facility is output; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5.
[0034] In one embodiment, the YOLOv5 network structure includes a backbone feature extraction network, a feature pyramid network, and a detection head network.
[0035] In a specific embodiment, the backbone feature extraction network (Backbone)—CSPDarknet: CSPDarknet is the core feature extraction module of YOLOv5, which is improved based on Darknet53 in YOLOv3 and adopts the Cross Stage Partial (CSP) strategy. The CSP strategy divides the feature map into two parts, one part is directly passed to the next layer, and the other part is spliced with the previous part after the convolution operation. This design effectively reduces the redundant information of the feature map and improves the computational efficiency of the model. CSPDarknet consists of multiple CSP residual blocks (CSPResBlock), each of which includes a CSP bottleneck (CSPBottleneck) and a residual block (Residual Block). The CSP bottleneck realizes the segmentation and splicing of the feature map, while the residual block enhances the feature extraction capability of the model by increasing the depth and nonlinear capability of the feature map. In addition, CSPDarknet also includes a focus module (Focus Module), which effectively improves the recognition ability of the model by reducing the image scale and increasing the number of channels, especially when facing the complex background of railway images, showing strong adaptability.
[0036] Feature Pyramid Network (Neck)—PANet: PANet is a feature pyramid network in YOLOv5, responsible for processing the prediction of multi-scale targets. PANet consists of a spatial pyramid pooling (SPP) module and a new CSP-PAN structure. The SPP module captures multi-scale features through maximum pooling layers of different sizes, enhancing the robustness of the features. The new CSP-PAN structure uses bottom-up and top-down feature fusion methods to balance the semantic information and spatial information of the feature map, thereby improving the quality of the feature map. In addition, by combining the CSP structure and residual blocks, PANet effectively reduces feature redundancy and increases model depth, providing strong support for the detection of complex morphological facilities.
[0037] Detection head network (Head) - Yolo Layer: Yolo Layer is the detection head module of YOLOv5, which is responsible for the final recognition and positioning of the target. This module optimizes the output layer so that the model can adapt to anchors of different sizes and proportions. Yolo Layer predicts four attributes for each anchor: center coordinates, width and height, confidence, and category probability. The center coordinates and width and height attributes are offsets relative to the anchor point, the confidence indicates the possibility of the existence of the target, and the category probability indicates the probability that the target belongs to each category. Through this structure, YOLOv5 can achieve efficient and accurate detection and classification when dealing with the diversity and complexity of facilities.
[0038] In order to further improve the recognition accuracy of the YOLOv5 network for diverse facility forms in complex backgrounds, the present invention integrates the Squeeze-and-Excitation structure in the residual network of the YOLOv5 network. The SE structure is a streamlined and efficient channel attention mechanism that dynamically adjusts the channel weights of the feature map to enhance the model's attention to key information, thereby improving feature expression capabilities and detection performance.
[0039] The SE structure aims to optimize the expressiveness of feature maps by introducing dependencies between channels. Specifically, SE achieves this goal through the following steps:
[0040] Global information aggregation: First, the SE module performs global average pooling (GAP) on the input feature map to compress the information of the spatial dimension into a global representation of each channel. This step obtains global spatial information by calculating the global mean of each channel and captures the overall characteristics of each channel in the feature map.
[0041] Modeling inter-channel dependencies: Subsequently, the aggregated global information is processed through a two-layer fully connected network (FC). The first layer of the FC reduces the channel dimension, thereby reducing the computational complexity and introducing nonlinear mapping to enhance the expressiveness of features; the second layer of the FC restores the channel dimension to its original size to generate the attention weight of each channel. This weight vector represents the importance of each channel in the current feature map.
[0042] Channel reweighting: Finally, the generated attention weights are applied to the original feature map through a channel-by-channel element-wise multiplication operation to reweight the features of each channel. In this way, the network can pay more attention to channels with high weights and appropriately suppress channels with low weights, thereby highlighting the features that are important to the task in the feature map and weakening the impact of interference information.
[0043] After integrating the SE structure into the YOLOv5 backbone network, the model's feature extraction capability has been further enhanced, especially when dealing with complex facility scenes. The addition of the SE module enables YOLOv5 to effectively enhance the detection of subtle and irregular facilities. In addition, since the SE structure achieves feature enhancement through a lightweight attention mechanism, the computational overhead remains within an acceptable range while improving the model's detection accuracy. This balance enables the improved YOLOv5 to show greater robustness and adaptability in facility positioning tasks, and to more accurately identify and locate different types of facility areas.
[0044] Figure 2This is a flow chart of establishing a line facility image target detection model in an embodiment of the present invention. In one embodiment, the line facility image target detection model is established based on the optimized YOLOv5 network structure, and further includes:
[0045] Step 201, inputting the line facility image into the backbone feature extraction network, sampling the input line facility image multiple times through the backbone feature extraction network, and outputting multiple feature maps of different scales;
[0046] Step 202, inputting multiple feature maps of different scales into a feature pyramid network, performing feature fusion on the multiple feature maps of different scales through the feature pyramid network, and obtaining a line facility image target detection model based on YOLOv5 multi-scale fusion.
[0047] In the specific embodiment, according to the information characteristics of railway line facilities, the YOLOv5-SENet-EIoU network is designed to accurately locate line facilities. The YOLOv5 network achieves effective detection of facilities of different sizes and shapes through multi-scale feature fusion and anchor point adaptation strategy, and has high detection speed and accuracy.
[0048] In one embodiment, the EIoU loss function is used to optimize the original loss function of YOLOv5, including:
[0049] According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined.
[0050] In one embodiment, the total loss function of the optimized YOLOv5 is determined according to the EIoU loss function, the cross entropy loss function, and the binary cross entropy loss function, including:
[0051] The loss value formula based on the EIoU loss function is as follows:
[0052]
[0053] Among them, L loc represents the EIoU loss function, IoU represents the intersection over union ratio, ρ(b, b gt ) represents the predicted bounding box center b and the true bounding box center b gt The Euclidean distance between them, Δw and Δh are the difference in width and height between the predicted bounding box and the true bounding box, respectively. gt and h gt are the width and height of the minimum bounding box covering the two predicted bounding boxes and the true bounding box respectively;
[0054] YOLO v The total loss function of 5 is expressed as:
[0055] L total =λ cls L cls +λ obj L obj +λ loc L loc
[0056] Among them, L cls represents the cross entropy loss function, λ cls is the corresponding weight coefficient; L obj represents the binary cross entropy loss function, λ obj is the corresponding weight coefficient; L loc represents the EIoU loss function, λ loc is the corresponding weight coefficient.
[0057] In a specific embodiment, the difference between the category probability predicted by the model and the actual label is measured by the category loss (Categorical Loss), and the category loss value is generally calculated by the cross entropy loss function; the objectness loss (Objectness Loss) is used to determine whether the anchor point contains an object, and the binary cross entropy loss function is generally used to calculate the object loss value; the difference between the predicted bounding box and the actual bounding box is measured by the localization loss (Localization Loss), and the mean square error loss function is generally used to calculate the location loss value. In order to improve the network positioning accuracy, the present invention uses the EIoU loss function to decouple the aspect ratio error of the bounding box from the center point error to improve the optimization efficiency.
[0058] In step 103, within the target detection area, OCR technology is used to extract text information from the line facility image, and the inventory data of the line facility is determined based on the text information.
[0059] In a specific embodiment, OCR technology is used to identify and extract text information from line facility images within the target detection area, extracting key information including line facility number, category, standard mileage, etc., thereby improving the accuracy of information extraction, ensuring the consistency and relevance of line facilities and their ledger data, and providing efficient and reliable technical support for the automated construction of electronic ledgers.
[0060] In a specific embodiment, the SVTR-tiny deep network may also be used to identify and extract text information of line facility images within the target detection area.
[0061] In step 104, the ledger data is uniformly stored in a pre-built database.
[0062] Figure 3This is a flow chart of obtaining ledger data through line facility identification in an embodiment of the present invention. In one embodiment, the ledger data is uniformly stored in a pre-built database, and further includes:
[0063] Step 301, deploying a unique identifier on a line facility;
[0064] Step 302, using a line array camera to identify the identification of the line facility, and obtaining the corresponding line facility ledger data; the line facility ledger data includes the line facility number, line facility category, and line facility mileage information.
[0065] In a specific embodiment, the identification information obtained by using a linear array camera is matched with the inventory data of the line facilities to establish a standardized and structured electronic equipment ledger, covering detailed contents such as equipment location, category, number, mileage, etc., so as to realize rapid query of line facility inventory data and provide a reliable data basis for subsequent analysis, query and maintenance.
[0066] The present invention also provides a line facility ledger data processing device, as described in the following embodiments. Since the principle of the device to solve the problem is similar to the line facility ledger data processing method, the implementation of the device can refer to the implementation of the line facility ledger data processing method, and the repeated parts will not be repeated.
[0067] Figure 4 Schematic diagram of a device for processing account data of line facilities in an embodiment of the present invention, the device comprising:
[0068] The line facility image acquisition module 401 is used to acquire the line facility image collected by the line array camera in the target line;
[0069] The target detection area determination module 402 is used to input the line facility image into the line facility image target detection model and output the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5;
[0070] The ledger data determination module 403 is used to extract text information of the line facility image using OCR technology in the target detection area, and determine the ledger data of the line facility according to the text information;
[0071] The ledger data storage module 404 is used to uniformly store the ledger data in a pre-built database.
[0072] In one embodiment, the YOLOv5 network structure includes a backbone feature extraction network, a feature pyramid network, and a detection head network.
[0073] In one embodiment, a line facility image target detection model building module is further included, which is used to:
[0074] Inputting the line facility image into the backbone feature extraction network, sampling the input line facility image multiple times through the backbone feature extraction network, and outputting multiple feature maps of different scales;
[0075] Multiple feature maps of different scales are input into the feature pyramid network, and the feature maps of multiple scales are fused through the feature pyramid network to obtain a line facility image target detection model based on YOLOv5 multi-scale fusion.
[0076] In one embodiment, the target detection area determination module 402 is further configured to:
[0077] According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined.
[0078] In one embodiment, the target detection area determination module 402 is specifically configured to:
[0079] According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined;
[0080] The loss value formula based on the EIoU loss function is as follows:
[0081]
[0082] Among them, L loc represents the EIoU loss function, IoU represents the intersection over union ratio, ρ(b, b gt ) represents the predicted bounding box center b and the true bounding box center b gt The Euclidean distance between them, Δw and Δh are the difference in width and height between the predicted bounding box and the true bounding box, respectively. gt and h gt are the width and height of the minimum bounding box covering the two predicted bounding boxes and the true bounding box respectively;
[0083] YOLO v The total loss function of 5 is expressed as:
[0084] L total =λ cls L cls +λ obj L obj +λ locL loc
[0085] Among them, L cls represents the cross entropy loss function, λ cls is the corresponding weight coefficient; L obj represents the binary cross entropy loss function, λ obj is the corresponding weight coefficient; L loc represents the EIoU loss function, λ loc is the corresponding weight coefficient.
[0086] In one embodiment, the ledger data storage module 404 is also used to:
[0087] Deploy unique identification on line facilities;
[0088] The line facility identification is identified by a line array camera to obtain the corresponding line facility ledger data; the line facility ledger data includes the line facility number, line facility category, and line facility mileage information.
[0089] An embodiment of the present invention further provides a computer device, Figure 5 It is a schematic diagram of a computer device in an embodiment of the present invention. The computer device 500 includes a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 830, the above-mentioned line facility ledger data processing method is implemented.
[0090] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned line facility ledger data processing method.
[0091] An embodiment of the present invention also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned line facility ledger data processing method.
[0092] In the embodiment of the present invention, the line facility image captured by the linear array camera is obtained in the target line; the line facility image is input into the line facility image target detection model, and the target detection area of the line facility is output; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the original loss function of YOLOv5 is optimized by the EIoU loss function; in the target detection area, the text information of the line facility image is extracted by the OCR technology, and the account data of the line facility is determined according to the text information; the account data is uniformly stored in a pre-built database. In the above process, the embodiment of the present invention is based on the optimized target detection YOLOv5 model to obtain the target detection area of the line facility, improve the accuracy of the positioning of the target detection area, and use the OCR technology to determine the text information of the line facility image in the target detection area as the account data, thereby improving the accuracy of the account data, and providing data support for the real-time management and mileage correction of the line facilities.
[0093] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0094] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0095] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.
[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0097] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for processing line facility ledger data, characterized in that: include: In the target line, obtain line facility images collected by the line array camera; Inputting the line facility image into the line facility image target detection model, and outputting the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5; In the target detection area, OCR technology is used to extract text information from line facility images, and the record data of line facilities is determined based on the text information; The ledger data is uniformly stored in a pre-built database.
2. The method according to claim 1, characterized in that The YOLOv5 network structure includes a backbone feature extraction network, a feature pyramid network, and a detection head network.
3. The method according to claim 2, characterized in that A line facility image target detection model is established based on the optimized YOLOv5 network structure, which also includes: Inputting the line facility image into the backbone feature extraction network, sampling the input line facility image multiple times through the backbone feature extraction network, and outputting multiple feature maps of different scales; Multiple feature maps of different scales are input into the feature pyramid network, and the feature maps of multiple scales are fused through the feature pyramid network to obtain a line facility image target detection model based on YOLOv5 multi-scale fusion.
4. The method according to claim 1, characterized in that The EIoU loss function is used to optimize the original loss function of YOLOv5, including: According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined.
5. The method according to claim 4, characterized in that According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined, including: The loss value formula based on the EIoU loss function is as follows: Among them, L loc represents the EIoU loss function, IoU represents the intersection over union ratio, ρ(b, b gt ) represents the predicted bounding box center b and the true bounding box center b gt The Euclidean distance between them, Δw and Δh are the difference in width and height between the predicted bounding box and the true bounding box, respectively. gt and h gt are the width and height of the minimum bounding box covering the two predicted bounding boxes and the true bounding box respectively; The total loss function of YOLOv5 is expressed as: L total =λ cls L cls +λ obj L obj +λ loc L loc Among them, L cls represents the cross entropy loss function, λ cls is the corresponding weight coefficient; L obj represents the binary cross entropy loss function, λ obj is the corresponding weight coefficient; L loc represents the EIoU loss function, λ loc is the corresponding weight coefficient.
6. The method according to claim 1, characterized in that The ledger data is uniformly stored in a pre-built database, which also includes: Deploy unique identification on line facilities; The line facility identification is identified by a line array camera to obtain the corresponding line facility ledger data; the line facility ledger data includes the line facility number, line facility category, and line facility mileage information.
7. A line facility ledger data processing device, characterized in that: include: A line facility image acquisition module is used to acquire line facility images captured by a linear array camera in a target line; The target detection area determination module is used to input the line facility image into the line facility image target detection model and output the target detection area of the line facility; the line facility image target detection model is established based on the optimized YOLOv5 network structure; in the process of optimizing the YOLOv5 network structure, the SENet network is embedded in the residual module of YOLOv5, and the EIoU loss function is used to optimize the original loss function of YOLOv5; The ledger data determination module is used to extract text information from line facility images using OCR technology within the target detection area, and determine the ledger data of the line facilities based on the text information; The ledger data storage module is used to uniformly store ledger data in a pre-built database.
8. The device according to claim 7, characterized in that The YOLOv5 network structure includes a backbone feature extraction network, a feature pyramid network, and a detection head network.
9. The device according to claim 8, characterized in that It also includes a line facility image target detection model building module for: Inputting the line facility image into the backbone feature extraction network, sampling the input line facility image multiple times through the backbone feature extraction network, and outputting multiple feature maps of different scales; Multiple feature maps of different scales are input into the feature pyramid network, and the feature maps of multiple scales are fused through the feature pyramid network to obtain a line facility image target detection model based on YOLOv5 multi-scale fusion.
10. The device according to claim 7, characterized in that The target detection area determination module is also used to: According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined.
11. The device according to claim 10, characterized in that The target detection area determination module is specifically used for: According to the EIoU loss function, cross entropy loss function and binary cross entropy loss function, the total loss function of the optimized YOLOv5 is determined; The loss value formula based on the EIoU loss function is as follows: Among them, L loc represents the EIoU loss function, IoU represents the intersection over union ratio, ρ(b, b gt ) represents the predicted bounding box center b and the true bounding box center b gt The Euclidean distance between them, Δw and Δh are the difference in width and height between the predicted bounding box and the true bounding box, respectively. gt and h gt are the width and height of the minimum bounding box covering the two predicted bounding boxes and the true bounding box respectively; The total loss function of YOLOv5 is expressed as: L total =λ cls L cls +λ obj L obj +λ loc L loc Among them, L cls represents the cross entropy loss function, λ cls is the corresponding weight coefficient; L obj represents the binary cross entropy loss function, λ obj is the corresponding weight coefficient; L loc represents the EIoU loss function, λ loc is the corresponding weight coefficient.
12. The device according to claim 7, characterized in that The ledger data storage module is also used for: Deploy unique identification on line facilities; The line facility identification is identified by a line array camera to obtain the corresponding line facility ledger data; the line facility ledger data includes the line facility number, line facility category, and line facility mileage information.
13. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Line electronic standing book defect positioning method and device
CN121637085A