Machine vision-based rail transit station obstacle monitoring method and device

By using lightweight deep learning models and multi-surveillance video fusion algorithms in rail transit stations, the problems of low efficiency and low accuracy of obstacle recognition in the prior art are solved, and higher recognition accuracy and reliability are achieved.

CN119992476APending Publication Date: 2025-05-13ZHEJIANG RAIL TRANSIT OPERATION MANAGEMENT GROUP CO LTD +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510122303.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing machine vision system has low efficiency and low accuracy in identifying obstacles in rail transit stations, and the overlapping coverage of multiple cameras leads to multiple identification and false alarms.

Method used

The lightweight deep learning model is used to combine the multi-surveillance video fusion algorithm to predict obstacle recognition results through the lightweight deep learning model, and the multi-surveillance video fusion algorithm is used to verify and deduplicate to determine the target obstacle.

Benefits of technology

It improves the accuracy and reliability of obstacle recognition, reduces false alarms and repeated recognition, and improves the accuracy and efficiency of the monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992476A_ABST
    Figure CN119992476A_ABST
Patent Text Reader

Abstract

The invention provides a rail transit station obstacle monitoring method and device based on machine vision, and relates to the technical field of machine vision and rail transit, and the method comprises the steps: inputting a to-be-recognized picture generated based on a monitoring video into a lightweight deep learning model, and predicting and generating an obstacle recognition result; and then based on an obstacle recognition result, a multi-monitoring video fusion algorithm is used for checking and determining a target obstacle, and the multi-monitoring video fusion algorithm is used for matching the recognized obstacle with monitoring videos of different video sources in the rail transit station and de-weighting the obstacle. Through the method, the technical problems of low station obstacle identification efficiency and low accuracy in the prior art are solved, and the technical effect of improving the station obstacle identification precision and reliability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine vision and rail transit technology, and in particular to a method and device for monitoring obstacles in rail transit stations based on machine vision. Background Art

[0002] Rail transit stations are typical crowded places, equipped with ticket gates, security inspection machines, ticket machines, vertical billboards and other facilities. In the event of an emergency, these facilities may interfere with the escape path of pedestrians. If these obstacles can be sensed and avoided in real time, thereby reducing their impact on pedestrian evacuation, this will play a significant role in the safe evacuation of high-density people in emergencies and help ensure the safe operation of rail transit.

[0003] In recent years, the rail transit industry has increasingly adopted machine vision technology to improve safety and operational efficiency. For example, machine vision technology can be used to monitor passenger behavior and identify abnormal behavior; in emergency evacuation, machine vision can dynamically plan the optimal evacuation route based on real-time passenger flow; machine vision can also be used to monitor the status of station equipment and detect track wear to prevent safety accidents. The application of machine vision technology has significantly improved the safety and operational efficiency of rail transit and enhanced the passenger experience. Despite this, the existing machine vision system still needs to be further optimized.

[0004] At present, the machine vision system needs to process a large amount of real-time video streaming data, which requires high computing power. The system may delay alarms due to slow processing speed, affecting the timeliness of emergency response; and the accuracy of machine vision depends largely on algorithms. Some algorithms perform poorly in complex scenes and have poor accuracy in distinguishing obstacles from backgrounds. In addition, due to the overlapping coverage of multiple surveillance cameras installed inside rail transit stations, the same obstacle may be captured by multiple cameras, resulting in multiple recognition and alarms of the same obstacle by the monitoring system, which not only increases the workload of monitoring personnel, but may also cause false alarms, affecting the accuracy and efficiency of the monitoring system. Summary of the invention

[0005] The purpose of the present invention is to provide a method and device for monitoring obstacles in rail transit stations based on machine vision, so as to alleviate the technical problems of low recognition efficiency and low accuracy existing in the prior art.

[0006] In the first aspect, an embodiment of the present invention provides a method for monitoring obstacles in rail transit stations based on machine vision, including: inputting a picture to be identified generated based on a surveillance video into a lightweight deep learning model, and predicting and generating obstacle identification results; based on the above obstacle identification results, using a multi-surveillance video fusion algorithm to perform inspection and determine the target obstacle; the above multi-surveillance video fusion algorithm is used to match the identified obstacles with surveillance videos from different video sources in the rail transit station, and deduplicate the above obstacles.

[0007] In some optional implementations, the method further includes: collecting surveillance videos through video sources at different locations in a rail transit station; generating pictures to be identified based on the surveillance videos, and preprocessing the pictures to be identified to generate feature maps to be identified; the step of inputting the pictures to be identified generated based on the surveillance videos into a lightweight deep learning model includes: inputting the preprocessed feature maps to be identified into the lightweight deep learning model.

[0008] In some optional implementations, the above-mentioned lightweight deep learning model includes: a backbone network, a feature enhancement network, a detection head and an output module; the above-mentioned backbone network includes a C2f module and an SPPF module; the above-mentioned C2f module is used to divide the above-mentioned feature map to be identified into four blocks according to RGBA information for separate processing; the above-mentioned SPPF module is used to extract multi-scale feature maps through multi-layer non-expanded convolution operations; the above-mentioned feature enhancement network is a dual-stream FPN structure, which is used to perform feature fusion on the above-mentioned multi-scale feature map; the above-mentioned detection head is used to make predictions based on the fused feature map to generate obstacle recognition results; the above-mentioned output module is used to output obstacle recognition results, and the above-mentioned obstacle recognition results include: target category, bounding box position and confidence score.

[0009] In some optional implementations, the dual-stream FPN structure of the feature enhancement network includes: a feature pyramid network and a path aggregation network, and the feature pyramid network and the path aggregation network are respectively used to fuse feature maps of different levels.

[0010] In some optional implementations, based on the above obstacle recognition results, a multi-surveillance video fusion algorithm is used to perform inspection and determine the target obstacle, including: based on the time information of the above surveillance video, verifying the time of the model recognition result, and aligning the frames of the same time points of the above surveillance videos obtained from different video sources; the above time information includes the timestamp and time frame of the current video; constructing a real three-dimensional space based on surveillance videos obtained from video sources at different locations in the rail transit station; performing feature point matching on the image information of the above surveillance video and the real three-dimensional space information, and saving the matching results in a feature point database.

[0011] In some optional implementations, based on the above obstacle recognition results, a multi-surveillance video fusion algorithm is used to perform inspection and determine the target obstacle, which also includes: matching the pixels of the above image information with the pixels of the above real three-dimensional space according to the matched feature points; based on the pixel matching results, performing coordinate transformation on the above obstacle to determine the space where the above obstacle exists; according to the category of the above obstacle and the space where the above obstacle exists, merging and deduplicating the same obstacles to determine the target obstacle.

[0012] In some optional implementations, the categories of the obstacles include color and shape; the space where the obstacles exist includes location and scene.

[0013] In the second aspect, an embodiment of the present invention provides a rail transit station obstacle monitoring device based on machine vision, which includes: an identification module, which is used to input a picture to be identified generated based on a surveillance video into a lightweight deep learning model, and predict and generate an obstacle identification result; a verification module, which is used to use a multi-surveillance video fusion algorithm to perform a verification based on the above obstacle identification result to determine the target obstacle; the above-mentioned multi-surveillance video fusion algorithm is used to match the identified obstacles with surveillance videos of different video sources in the rail transit station, and deduplicate the above-mentioned obstacles.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the steps of any one of the methods described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute any method described in the first aspect.

[0016] The present invention provides a method and device for monitoring obstacles in rail transit stations based on machine vision, the method comprising: inputting a picture to be identified generated based on a surveillance video into a lightweight deep learning model, predicting and generating obstacle identification results; then based on the obstacle identification results, using a multi-surveillance video fusion algorithm for inspection to determine the target obstacle, the multi-surveillance video fusion algorithm is used to match the identified obstacle with surveillance videos of different video sources in the rail transit station, and deduplicate the obstacle. The above method solves the technical problems of low efficiency and low accuracy of station obstacle identification in the prior art, and achieves the technical effect of improving the accuracy and reliability of station obstacle identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A schematic diagram of a flow chart of a method for monitoring obstacles in a rail transit station based on machine vision provided by an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of the structure and application of a lightweight deep learning model and a multi-surveillance video fusion algorithm provided in an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of the structure of a rail transit station obstacle monitoring device based on machine vision provided by an embodiment of the present invention;

[0021] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0023] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. Some embodiments of the present invention are described in detail below in conjunction with the drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0025] Machine vision technology is gradually being widely used in the field of rail transit to improve rail transit safety and operational efficiency. For example, machine vision technology can be used to analyze video streams and images to assist the fire alarm system (FAS) in early detection of fires and identify smoke and flame characteristics. Machine vision technology can also be used to monitor passenger behavior, identify abnormal behaviors such as running or fighting, and notify security personnel in a timely manner. In emergency evacuation, machine vision can dynamically plan the optimal evacuation route based on real-time passenger flow. In addition, machine vision can also be used to monitor the status of station equipment, such as elevators and escalators, as well as detect track wear and prevent safety accidents. In terms of service, machine vision technology is integrated into self-service machines to provide fast identity authentication and personalized services through facial recognition. For visually impaired passengers, machine vision assists in providing navigation services. Machine vision also collects passenger behavior data to help optimize service processes. In short, the application of machine vision makes rail transit safer and more efficient, and improves the passenger experience.

[0026] However, the machine vision system needs to process a large amount of real-time video streaming data, which requires high computing power. The system may delay alarms due to slow processing speed, affecting the timeliness of emergency response. In addition, the accuracy of machine vision depends largely on algorithms. Some algorithms perform poorly in complex scenarios and have poor accuracy in distinguishing obstacles from backgrounds. In addition, due to the overlapping coverage of multiple surveillance cameras installed inside rail transit stations, the same obstacle may be captured by multiple cameras, resulting in multiple recognition and alarms of the same obstacle by the monitoring system. This not only increases the workload of monitoring personnel, but may also lead to false alarms, affecting the accuracy and efficiency of the monitoring system.

[0027] Based on this, an embodiment of the present invention provides a method and device for monitoring obstacles in rail transit stations based on machine vision to solve the technical problems of low efficiency and low accuracy in identifying obstacles in stations in the prior art.

[0028] To facilitate understanding of this embodiment, a method for monitoring obstacles in a rail transit station based on machine vision disclosed in an embodiment of the present invention is first introduced in detail. Figure 1 The flowchart of a method for monitoring obstacles in a rail transit station based on machine vision is shown. The method can be executed by an electronic device and mainly includes the following steps S110 to S120:

[0029] S110: Inputting the image to be identified generated based on the surveillance video into the lightweight deep learning model to predict and generate obstacle identification results;

[0030] The machine vision system needs to process a large amount of real-time video stream data, which places high demands on computing power. In the case of huge amounts of data, the system may delay alarms because the processing speed cannot keep up, affecting the timeliness of emergency response. In addition, the accuracy of machine vision depends largely on algorithms. Some algorithms perform poorly in complex scenarios (such as obstacles of different shapes, sizes, and colors), and it is difficult to accurately distinguish obstacles from the background. In this regard, this embodiment improves the model structure based on the principles of machine vision algorithms to achieve the purpose of reducing the amount of calculation.

[0031] In one embodiment, the lightweight deep learning model includes: a backbone network, a feature enhancement network, a detection head and an output module; the backbone network includes a C2f module and an SPPF module; the C2f module is used to divide the feature map to be identified into four blocks for separate processing according to the RGBA information; the SPPF module is used to extract multi-scale feature maps through multi-layer non-expanded convolution operations. The feature enhancement network is a two-stream FPN structure, which is used to perform feature fusion on the multi-scale feature map; the detection head is used to make predictions based on the fused feature map to generate obstacle recognition results. The output module is used to output the obstacle recognition results, and the obstacle recognition results include: target category, bounding box position and confidence score.

[0032] In one embodiment, the dual-stream FPN structure of the feature enhancement network includes: a feature pyramid network and a path aggregation network, and the feature pyramid network and the path aggregation network are respectively used to fuse feature maps of different levels.

[0033] That is to say, in view of the shortcomings and deficiencies of existing machine vision methods, the above optimization method is proposed, that is, a lighter network architecture (the above lightweight deep learning model) is adopted, and the optimized C2f module and SPPF module are adopted in the backbone. In the C2f module, the input feature map is divided into four blocks according to the RGBA color in the initial dimension, which allows the model to process four feature maps separately, improves the parallelism and computational efficiency of the model, and the number of channels of the input Tensor of each BottleNeck is only 0.25 times that of the previous level, so the amount of calculation is significantly reduced, which improves the efficiency and accuracy of feature extraction.

[0034] In the SPPF module, non-expanded convolution is used to replace the traditional convolution operation, thereby expanding the receptive field without increasing the amount of calculation; secondly, the advanced loss function Distribution Focal Loss (DFL) and the positive sample assignment strategy TaskAlignedAssigner are introduced. For each prediction sequence, DFL considers the distance between the prediction box and the label box, and calculates the loss of the left and right boundaries. In this way, DFL can more effectively focus on those difficult-to-classify samples. TaskAlignedAssigner combines DFL to optimize sample assignment and loss calculation. It increases the flexibility of positive sample box selection through a dynamic label matching strategy. This means that the model can more flexibly adjust the selection of positive samples according to the current prediction situation, so as to better adapt to unbalanced data sets and enhance the performance of the model when processing unbalanced data sets; in addition, a two-stream FPN structure is adopted to effectively integrate multi-scale features and improve the detection ability of small targets; finally, the Anchor-Free idea is adopted to directly predict the target center point and the bounding box regression, which further simplifies the model structure and improves the generalization ability and adaptability of the model.

[0035] In one embodiment, before the above-mentioned step S110, the method may also include the acquisition and processing of surveillance videos, namely: acquiring surveillance videos through video sources at different locations in the rail transit station; then generating a picture to be identified based on the surveillance video, and preprocessing the picture to be identified to generate a feature map to be identified; and then inputting the preprocessed feature map to be identified into a lightweight deep learning model.

[0036] Among them, the video source used to collect surveillance videos is generally surveillance cameras installed at different locations and angles in rail transit stations.

[0037] The embodiments of the present invention introduce lightweight deep learning models that reduce the number of model parameters and computational complexity while maintaining a high recognition accuracy, allowing the algorithm to run faster on edge devices. Secondly, an attention mechanism is introduced to enable the algorithm to automatically focus on key areas in the image, thereby improving the accuracy of obstacle recognition. In this way, the algorithm can allocate computing resources more efficiently, focusing on areas that are most likely to contain obstacles, rather than processing the entire image uniformly.

[0038] S120: Based on the obstacle recognition result, a multi-surveillance video fusion algorithm is used to perform inspection and determine the target obstacle; the multi-surveillance video fusion algorithm is used to match the identified obstacle with the surveillance videos of different video sources in the rail transit station, and deduplicate the obstacles.

[0039] Usually, multiple surveillance cameras are installed inside rail transit stations to achieve comprehensive monitoring of every corner of the station. However, due to the overlap of camera coverage, the same obstacle may be captured by multiple cameras, resulting in multiple recognition and alarms of the same obstacle by the monitoring system, which not only increases the workload of monitoring personnel, but may also cause false alarms, affecting the accuracy and efficiency of the monitoring system.

[0040] In order to solve the above problems, this embodiment proposes a multi-surveillance video fusion algorithm, the core idea of ​​which is to achieve accurate identification and positioning of obstacles by intelligently analyzing and processing video data captured by multiple cameras.

[0041] In one embodiment, S120 uses a multi-monitoring video fusion algorithm to perform inspection based on the obstacle recognition result to determine the target obstacle, including:

[0042] (S21) based on the time information of the surveillance video, verifying the time of the model recognition result, and aligning the frames of the surveillance videos obtained from different video sources at the same time point; the time information includes the timestamp and time frame of the current video;

[0043] (S22) constructing a real three-dimensional space based on surveillance videos acquired from video sources at different locations in the rail transit station; performing feature point matching on image information of the surveillance videos and real three-dimensional space information, and storing the matching results in a feature point database.

[0044] In one embodiment, based on the obstacle recognition result, a multi-monitoring video fusion algorithm is used to perform inspection and determine the target obstacle, further comprising:

[0045] (S23) matching pixels of the image information with pixels of the real three-dimensional space according to the matched feature points;

[0046] (S24) Based on the pixel matching result, coordinate transformation is performed on the obstacle to determine the space where the obstacle exists;

[0047] (S25) According to the types of obstacles and the spaces where the obstacles exist, identical obstacles are merged and deduplicated to determine the target obstacle.

[0048] In one embodiment, the categories of obstacles include color and shape; the space where the obstacles exist includes: position and scene.

[0049] That is to say, the embodiment of the present invention proposes a new multi-surveillance video fusion algorithm, which ensures that all input video streams are synchronized based on a unified time standard through precise time synchronization and alignment, and analyzes and adjusts the video frame rate, timestamp and other information to ensure that frames at the same time point in different video sources can be correctly aligned.

[0050] After time alignment, identify and match spatial feature points and image feature points. These feature points can be corners, edges or other significant features in the scene. By matching the feature points, the corresponding relationship between the position of the obstacle in the three-dimensional space and its position in the two-dimensional image is established. Next, the spatial coordinates of the obstacles identified in the video are transformed by the algorithm, which involves converting the obstacles from their original three-dimensional spatial coordinates to two-dimensional image coordinates. It is necessary to consider factors such as the camera's viewing angle and position to ensure that the converted coordinates can accurately reflect the actual position of the obstacle in the image. The spatial position of the obstacle is projected into a unified plane space to generate the plane coordinate value of the obstacle, so that the data in different video sources can be compared and analyzed on the same plane. By comparing the similarity of the plane coordinate values ​​of obstacles in different videos, including but not limited to the category characteristics and spatial characteristics of the obstacles, duplicate obstacles are identified and deduplication is performed.

[0051] The embodiment of the present invention achieves significant improvements in obstacle monitoring in rail transit stations by adopting an advanced multi-surveillance video fusion algorithm. The algorithm can integrate video streams captured by surveillance cameras from different locations and angles to comprehensively identify and analyze the same obstacle. In this way, the system can not only observe obstacles from multiple perspectives, but also use the data of each camera to calibrate each other, thereby reducing the recognition bias and error that may be caused by a single perspective. This fusion of multi-angle and multi-source data greatly improves the accuracy and reliability of obstacle identification.

[0052] As a specific example, combining Figure 2 As shown, this embodiment proposes a rail transit station obstacle monitoring method based on machine vision, which includes two parts: model recognition and obstacle monitoring.

[0053] Among them, the model recognition part is implemented through a lightweight deep learning model, including:

[0054] (1) The surveillance video is processed into images, and the images are resized to 640x640 pixels, normalized, and the pixel values ​​of the images are scaled to [0,1]. The images are then preprocessed by random cropping, rotation, inversion, and color adjustment to improve the generalization ability of the model.

[0055] (2) The preprocessed image is input into the backbone network, and the multi-scale feature map of the image is extracted through multi-layer convolution operations.

[0056] (3) Feature pyramid network (FPN) and path aggregation network (PAN) are used to fuse feature maps at different levels. Through cross-layer connection and feature fusion, the model's ability to monitor targets of different scales is improved.

[0057] In the application scene of rail transit stations, obstacles (such as railroad horses, ticket machines, security inspection machines, etc.) are generally large in size, and small obstacles generally do not affect the movement of pedestrians. Therefore, only large obstacles need to be detected. Therefore, in FPN, the scale range can be set larger, such as strengthening the P3 and P4 features of the feature pyramid, and using a higher resolution (512x512).

[0058] (4) A 1x1 convolutional layer is set on each detection head to generate predictions, including the object category, bounding box location, and confidence score.

[0059] The obstacle monitoring part can verify the recognition results based on the output results of the model using a multi-surveillance video fusion algorithm, including: (1) First, based on the timestamp and time frame of the video, the time of the model recognition result is verified to ensure that all video streams are synchronized based on a unified time standard. (2) The image information of the surveillance video and the real spatial information are matched with feature points and saved in the feature point database to achieve timely reuse. (3) According to the matched feature points, the pixels of the image information and the pixels of the real space are matched to realize the conversion of obstacles from image coordinates to spatial plane coordinates. (4) The same obstacles are merged and deduplicated according to the category of the obstacle (color, shape, etc.) and the space where the obstacle exists (position, scene, etc.).

[0060] That is to say, the embodiment of the present invention provides a method for monitoring obstacles in rail transit stations based on machine vision, which uses a lightweight deep learning model to identify obstacles in surveillance videos of different video sources in a rail transit station, and then uses a multi-surveillance video fusion algorithm to verify the recognition results based on the output results of the model, thereby determining the target obstacle. The lightweight deep learning model proposed in this embodiment can reduce the number of parameters and computational complexity of the model while maintaining a high recognition accuracy, and can more effectively allocate computing resources, focusing on the areas most likely to contain obstacles, so that the output recognition results are more accurate; the proposed multi-surveillance video fusion algorithm can integrate video streams captured by surveillance cameras from different positions and angles, and comprehensively identify and analyze the same obstacle. It can not only observe obstacles from multiple perspectives, but also use the data of each camera to calibrate each other, thereby reducing the recognition bias and error that may be caused by a single perspective, greatly improving the accuracy and reliability of obstacle recognition.

[0061] In addition, the embodiment of the present invention also provides a rail transit station obstacle monitoring device based on machine vision, combined with Figure 3 As shown, the device comprises:

[0062] The recognition module 310 is used to input the image to be recognized generated based on the surveillance video into the lightweight deep learning model to predict and generate obstacle recognition results;

[0063] The inspection module 320 is used to use a multi-surveillance video fusion algorithm to perform inspection based on the obstacle recognition result to determine the target obstacle; the multi-surveillance video fusion algorithm is used to match the identified obstacle with the surveillance videos of different video sources in the rail transit station, and to deduplicate the obstacles.

[0064] The machine vision-based rail transit station obstacle monitoring device provided in the embodiment of the present application can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiment of the present application has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brief description, for parts not mentioned in the device embodiment, reference can be made to the corresponding contents in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here. The machine vision-based rail transit station obstacle monitoring device provided in the embodiment of the present application has the same technical features as the machine vision-based rail transit station obstacle monitoring method provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0065] An embodiment of the present application further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above-mentioned embodiments.

[0066] Figure 4 A structural schematic diagram of an electronic device provided in an embodiment of the present application, the electronic device 400 includes: a processor 40, a memory 41, a bus 42 and a communication interface 43, wherein the processor 40, the communication interface 43 and the memory 41 are connected via the bus 42; the processor 40 is used to execute an executable module stored in the memory 41, such as a computer program.

[0067] The memory 41 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.

[0068] The bus 42 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0069] Among them, the memory 41 is used to store programs, and the processor 40 executes the program after receiving the execution instruction. The method executed by the process definition device disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 40 or implemented by the processor 40.

[0070] The processor 40 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 40. The above processor 40 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present invention can be directly embodied as a hardware decoding processor to execute, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 41, and the processor 40 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.

[0071] Corresponding to the above method, an embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the steps of the above method.

[0072] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0073] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0074] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0075] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0076] It should be noted that similar numbers and letters represent similar items in the accompanying drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for monitoring obstacles in rail transit stations based on machine vision, characterized in that: include: Input the image to be identified generated based on the surveillance video into the lightweight deep learning model to predict and generate obstacle recognition results; Based on the obstacle recognition result, a multi-monitoring video fusion algorithm is used to perform inspection and determine the target obstacle; The multi-surveillance video fusion algorithm is used to match the identified obstacles with the surveillance videos of different video sources in the rail transit station, and to deduplicate the obstacles.

2. The method according to claim 1, characterized in that The method further comprises: Collect surveillance videos through video sources at different locations in rail transit stations; Generate a picture to be identified based on the surveillance video, and pre-process the picture to be identified to generate a feature map to be identified; The steps of inputting the image to be identified generated based on the surveillance video into the lightweight deep learning model include: The preprocessed feature map to be identified is input into a lightweight deep learning model.

3. The method according to claim 2, characterized in that The lightweight deep learning model includes: a backbone network, a feature enhancement network, a detection head and an output module; The backbone network includes a C2f module and an SPPF module; the C2f module is used to divide the feature map to be identified into four blocks for separate processing according to RGBA information; the SPPF module is used to extract a multi-scale feature map through a multi-layer non-expanded convolution operation; The feature enhancement network is a dual-stream FPN structure, which is used to perform feature fusion on the multi-scale feature map; the detection head is used to make predictions based on the fused feature map to generate obstacle recognition results; The output module is used to output obstacle recognition results, and the obstacle recognition results include: target category, bounding box position and confidence score.

4. The method according to claim 3, characterized in that The dual-stream FPN structure of the feature enhancement network includes: a feature pyramid network and a path aggregation network, and the feature pyramid network and the path aggregation network are respectively used to fuse feature maps of different levels.

5. The method according to claim 1, characterized in that Based on the obstacle recognition result, a multi-monitoring video fusion algorithm is used to perform inspection and determine the target obstacle, including: Based on the time information of the surveillance video, the time of the model recognition result is verified, and the frames of the surveillance video obtained from different video sources at the same time point are aligned; the time information includes the timestamp and time frame of the current video; Based on the surveillance videos obtained from video sources at different locations in the rail transit station, a real three-dimensional space is constructed; feature point matching is performed on the image information of the surveillance video and the real three-dimensional space information, and the matching results are stored in a feature point database.

6. The method according to claim 5, characterized in that Based on the obstacle recognition result, a multi-monitoring video fusion algorithm is used to perform inspection to determine the target obstacle, and further includes: Matching pixels of the image information with pixels of the real three-dimensional space according to the matched feature points; Based on the pixel matching result, coordinate transformation is performed on the obstacle to determine the space where the obstacle exists; According to the categories of the obstacles and the spaces where the obstacles exist, identical obstacles are merged and deduplicated to determine target obstacles.

7. The method according to claim 6, characterized in that The categories of the obstacles include color and shape; The space where the obstacle exists includes: location and scene.

8. A rail transit station obstacle monitoring device based on machine vision, characterized in that: include: The recognition module is used to input the image to be recognized based on the surveillance video into the lightweight deep learning model to predict and generate obstacle recognition results; A detection module, used to detect and determine the target obstacle based on the obstacle recognition result using a multi-monitoring video fusion algorithm; The multi-surveillance video fusion algorithm is used to match the identified obstacles with the surveillance videos of different video sources in the rail transit station, and to deduplicate the obstacles.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Monitoring identification method and device based on machine vision

    CN121121632A

  • A monitoring and recognition method and device based on machine vision

    CN121121632B

  • High-precision field obstacle rapid detection method and system based on remote sensing image

    CN121353940A

  • A high-precision field obstacle rapid detection method and system based on remote sensing images

    CN121353940B