A real-time visual recognition and detection system for objects left in subway terminal carriages

Through the real-time visual recognition and detection system for objects left in subway terminal carriages, using image acquisition, feature extraction, time series association and anomaly judgment modules, rapid and accurate detection and risk assessment of objects left in subway carriages are achieved, solving the problems of low efficiency and insufficient accuracy in existing technologies and improving operational management efficiency.

CN120526231BActive Publication Date: 2025-09-23SHANGHAI BOZHIWEI ELECTRONIC SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511015168.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-09-23
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing technologies are inefficient and inaccurate in detecting debris in subway cars, making it difficult to achieve real-time detection. Furthermore, there is a lack of effective risk assessment of debris, resulting in inefficient operational management.

Method used

A real-time visual recognition and detection system for debris left in subway terminal carriages was designed. The image acquisition module generates carriage image sequences, the feature extraction module marks stable and changing areas, the temporal association module analyzes feature matching strength, the anomaly determination module identifies abnormal points, and the risk output module generates real-time detection and risk warning results.

Benefits of technology

It achieves rapid and accurate detection of objects left in carriages, reduces misjudgments and missed detections, can complete full-coverage detection during peak train traffic hours, and provide targeted risk warnings, thereby improving the accuracy of operational management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526231B_ABST
    Figure CN120526231B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of real-time detection systems for carriages of objects left behind, and discloses a real-time visual recognition and detection system for objects left behind in a subway terminal carriage. In this system, an image acquisition module acquires a surveillance video stream and generates a carriage image sequence set; a feature extraction module processes images to obtain a temporal correlation partition annotation set; a temporal correlation module analyzes temporal stable segment image frames to generate feature matching and superposition analysis results; an anomaly determination module identifies anomalies to form an anomaly point set for objects left behind; and a risk output module marks points with risk of objects left behind to generate real-time detection and risk warning results. The system can adapt to the dynamic environment of subway carriages, improve the efficiency and accuracy of object detection, and achieve real-time identification and risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time detection systems for carriage objects, and in particular to a real-time visual recognition and detection system for objects left in a carriage at a subway terminal. Background Art

[0002] The rapid development of urban rail transit networks has made the subway a core choice for daily commuting. With daily passenger volume continuing to climb, the complexity of the train environment has also increased. Passengers often leave items behind during boarding and alighting, either due to carelessness or emergencies. If these items are not promptly discovered and addressed, they can lead to a host of problems. For example, perishable items like food and beverages left unattended for extended periods can contaminate the train environment. Valuable items like electronic devices can cause financial loss. Unidentified packages and liquids can cause panic and even pose safety threats.

[0003] Currently, subway operators rely primarily on manual inspections to detect items left behind in terminal carriages. Upon arrival at the final station, staff must conduct a carriage-by-carriage inspection, a labor-intensive process that also limits inspection time. During peak hours, when train turnover is short, manual inspections often require only a quick glance, making it difficult to thoroughly examine hidden locations like under seats and in the corners of luggage racks. This can lead to some items being missed. Furthermore, manual judgment is susceptible to subjective factors, and different staff members may have different definitions of "left behind," potentially leading to misjudgments or omissions.

[0004] With the widespread adoption of video surveillance technology in subway systems, some lines are experimenting with using surveillance footage playback to aid in object detection. However, this method still relies on manual frame-by-frame review. The amount of video data captured by surveillance cameras is enormous; a single carriage's operating cycle can span hours of footage. Manual review is time-consuming, making real-time detection difficult. Furthermore, surveillance footage is susceptible to changes in lighting, such as the alternating light and dark when entering and exiting tunnels, and the turning on and off of lights in carriages at night. These factors can cause image quality fluctuations, complicating manual identification.

[0005] Existing computer vision-based detection solutions perform poorly in dynamic scenes. Traditional object detection algorithms are mostly designed for static images and cannot effectively handle dynamic changes within the train compartment. Passengers' movements such as standing up and walking around can cause temporary changes in the position of objects. Traditional algorithms may mistakenly identify items temporarily placed by passengers as left behind. When left behind are briefly obscured by other passengers and then reappear, they may be missed due to feature interruptions. Furthermore, the complex textures of fixed features within the train compartment, such as seats and armrests, share similarities with the features of some left behind, easily interfering with detection results.

[0006] When it comes to time series analysis, existing technologies lack the ability to deeply explore image sequences. The formation of artifacts exhibits distinct temporal characteristics: items remain stationary after a passenger leaves. However, existing algorithms often analyze only single frames, ignoring temporal variations in these characteristics. For example, an object may appear fixed across multiple frames and appear after the passenger's departure. This characteristic is crucial for identifying artifacts, but traditional algorithms struggle to capture this temporal correlation.

[0007] Furthermore, the existing system lacks effective risk assessment for left-behind items. Detection results only indicate "item present," failing to distinguish between a standard backpack and suspicious items, nor can the location of an item determine its impact on driving safety or passenger access. Upon receiving an alert, operators must conduct on-site verification of all alert points, reducing efficiency and potentially disrupting operational order through overreaction.

[0008] The existence of these problems has led to the fact that the detection of debris in subway cars has always faced difficulties such as low efficiency, insufficient accuracy, and delayed response, making it difficult to meet the high requirements of modern subway operations for safety and service. Summary of the Invention

[0009] The purpose of the present invention is to provide a real-time visual recognition and detection system for objects left in a subway terminal carriage to solve the problems raised in the above-mentioned background technology.

[0010] To achieve the above-mentioned object, the present invention provides a real-time visual recognition and detection system for objects left in a subway terminal carriage, the system comprising:

[0011] The image acquisition module obtains the surveillance video stream inside the terminal carriage, collects the carriage interior images of the corresponding time period, associates the time nodes and marks the image sequence index, and generates a carriage image sequence set;

[0012] The feature extraction module extracts image features and corresponding time indexes from the carriage image sequence set, arranges adjacent image pairs in chronological order, marks stable and changing regions in the adjacent image pairs, and obtains a temporal correlation partition annotation set;

[0013] The temporal association module obtains the image frames located in the temporally stable segment of the temporal association partition annotation set, extracts the pixel values ​​and color time series of the target area, compares the change amplitude with the common features of the remains, evaluates the feature matching strength under image fluctuations, and generates feature matching superposition analysis results;

[0014] The anomaly determination module identifies anomalies in the image frame whose feature matching value is greater than a reference matching threshold and is in a time sequence change section in the feature matching superposition analysis result, and forms a set of remnant anomaly points;

[0015] The risk output module obtains all points in the abnormal point set of the leftovers and the corresponding point information, marks the points with leftover risks, and generates real-time detection and risk warning results of leftovers in the carriage.

[0016] Preferably, the carriage image sequence set includes image feature values, image time indexes and normalized image factors; the time-series associated partition annotation set specifically includes time-series stable segment annotations, time-series changing segment annotations and feature difference rates of adjacent image frames; the feature matching superposition analysis results include the degree of influence of the pixel value change rate on the feature, the degree of influence of the color value change rate on the feature and the feature matching comparison under each image fluctuation condition; the residue anomaly point set includes the time position of the anomaly point, the color variation characteristics of the anomaly point pixel and the ratio of the anomaly point feature to the benchmark fluctuation; the real-time detection and risk warning results of residues in the carriage include a list of detected anomaly points and a joint judgment label of the three indicators of the anomaly point.

[0017] Preferably, the image acquisition module includes:

[0018] The video stream acquisition submodule acquires the surveillance video stream data in the terminal carriage, collects the time nodes and corresponding carriage interior images in the video stream, and records the acquisition results as two basic factors, time factor and image factor, to obtain the carriage image basic data set;

[0019] The image preprocessing submodule performs normalization processing based on the time factor and image factor data in the carriage image basic data group, establishes a corresponding relationship between the normalized results and the time nodes, calculates the average value of the normalized time value and the normalized image value as the image feature value, and generates a carriage image sequence set.

[0020] Preferably, the feature extraction module includes:

[0021] The feature extraction submodule obtains the image feature values ​​and corresponding time data in the carriage image sequence set, identifies the positional relationship of all image frames in the time sequence based on the time information, calls the image time index set, and uses the adjacent time threshold as a reference to perform time measurement and sort the image frames in the time sequence to generate a time sorted sequence of adjacent image frames;

[0022] The feature comparison submodule calculates the feature difference rate between the i-th and j-th adjacent image frames based on the time-ordered sequence of the adjacent image frames, and integrates them to generate a feature difference rate sequence;

[0023] The trend classification submodule extracts the pixel change trend and color change trend between adjacent image frames based on the feature difference rate sequence, classifies and labels each pair of image frames according to whether the two trend change directions are consistent, records and groups the segments with consistent trends and conflicting trends respectively, and obtains a time-series associated partition annotation set.

[0024] Preferably, the timing association module includes:

[0025] The sequence extraction submodule selects the segments marked as time-stable according to the time-series associated partition annotation set, detects the pixel values ​​and relative color value data of the target area within each image frame acquisition time period, arranges them in chronological order to form pixel time series and color time series, and generates an image change time series set;

[0026] The change rate calculation submodule calculates the pixel value change rate and color value change rate between consecutive time nodes in each image frame time series based on the image change time series set, compares the pixel value change rate and color value change rate in parallel under the same feature conditions, identifies the numerical relationship of the change amplitude under the feature by jointly analyzing the two types of rate indicators, integrates the influence value sequence of each image frame, and establishes the feature matching superposition analysis result.

[0027] Preferably, the abnormality determination module includes:

[0028] The record extraction submodule selects image frames whose feature matching values ​​are greater than the reference matching threshold and image frames in the time-series change section based on the feature matching overlay analysis results, extracts the continuous acquisition records of the image frames in chronological order, collects the pixel value and color value data of the target area corresponding to each time node, and generates a continuous acquisition record set;

[0029] The parameter change ratio calculation submodule calls the continuous acquisition record set, extracts the pixel values ​​and color values ​​of the image frame at two consecutive time nodes, calculates the pixel change ratio and color change ratio respectively, integrates them into a pixel change ratio sequence and a color change ratio sequence, and establishes an image fluctuation change data set;

[0030] The anomaly recognition submodule extracts the pixel distribution data and color distribution data of the corresponding time period based on the image fluctuation change data set, determines whether the pixel change ratio and color change ratio both exceed the set fluctuation recognition threshold, and determines whether the pixel distribution and color distribution deviate from the baseline distribution range at the same time. The time nodes that meet the conditions are marked as anomalies to generate a set of relic anomaly points.

[0031] Preferably, the risk output module includes:

[0032] The indicator joint judgment submodule obtains all points in the residue abnormal point set and the corresponding time index and position information, calculates the joint risk judgment value of the k-th image frame, and establishes a joint risk judgment value sequence;

[0033] Based on the joint risk judgment value sequence, the abnormal output sorting submodule screens points whose feature matching values ​​are greater than the feature matching risk threshold, image feature values ​​are lower than the feature reference value, and whose time series association labels are change segments, extracts the corresponding point time index, location identifier and partition to which it belongs, marks them as detected abnormalities with residual risks, and outputs the points that meet the joint conditions in a structured format, generating real-time detection and risk warning results for residual objects in the carriage.

[0034] Preferably, the video stream acquisition submodule includes acquiring video stream data from multiple surveillance cameras in the terminal car, synchronously calibrating the video stream to ensure the accurate correspondence between time nodes and image frames, collecting car interior images from different angles and recording them as multi-angle image factors, and acquiring a multi-view car image basic data group.

[0035] Preferably, the feature comparison submodule includes grayscale processing of adjacent image frames, extracting image edge contour features, calculating contour similarity as a basis for evaluating feature difference rate, and integrating to generate a feature difference rate sequence including contour similarity.

[0036] Preferably, the parameter change ratio calculation submodule includes calculating the ratio of the absolute difference between the pixel values ​​of two consecutive time nodes to the pixel value of the previous time node as the pixel change ratio, and calculating the ratio of the absolute difference between the color values ​​of two consecutive time nodes to the color value of the previous time node as the color change ratio, and integrating them into a pixel change ratio sequence and a color change ratio sequence containing the absolute difference ratio.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] This system is designed to meet the actual needs of detecting debris in carriages at subway terminal stations. Through the collaborative operation of multiple modules, it has formed a complete real-time identification solution, which effectively makes up for the many shortcomings of existing technologies. In terms of detection efficiency, the traditional manual inspection mode is limited by manpower allocation and work pace, making it difficult to achieve full-time and full-coverage detection. However, this system uses the image acquisition module to process the monitoring video stream in real time and automatically generate a set of carriage image sequences. Subsequent feature extraction, time series association and other modules can complete data processing without human intervention. The entire process from image acquisition to risk warning can be completed in a short time. This automated operation mode breaks the time limit of manual inspection. Even during peak hours when trains arrive at the station intensively, the detection of debris in each carriage can be completed quickly, avoiding the long-term detention of debris due to detection delays.

[0039] In terms of detection accuracy, existing technologies often make misjudgments due to their inability to distinguish between temporary items in dynamic changes and real relics. The feature extraction module of this system can clearly separate fixed facilities and moving targets in the carriage by marking stable and changing areas in adjacent image pairs, laying the foundation for subsequent identification. The temporal association module further focuses on the image frames in the temporal stable segment, extracts the pixel values ​​and color time series of the target area, and analyzes the amplitude of change in combination with common features of relics, effectively filtering out interference factors such as light fluctuations and lens shake, making the feature matching results more in line with the actual situation. The anomaly judgment module accurately locks on those anomalies that appear in the temporal change segment and have a high degree of feature matching by setting a baseline matching threshold, greatly reducing the situation where items temporarily placed by passengers are misjudged as relics, and also avoiding the problem of missing real relics due to feature interference.

[0040] Faced with the complex environmental conditions inside the carriage, this system demonstrates strong adaptability. Light changes, obstructions, and other conditions can cause image features to become unstable, making traditional detection methods susceptible to failure. However, this system's temporal correlation module generates overlay analysis results by evaluating the strength of feature matching under image fluctuations, enabling effective identification of features of objects left behind even when image quality fluctuates. For example, when the light inside the carriage suddenly dims, the system analyzes the changing patterns of the color time series to distinguish between overall tonal changes caused by the light and the characteristic differences of the objects themselves, ensuring that the recognition results are not affected by sudden changes in the environment.

[0041] In addition, the risk assessment design of this system is also superior to existing technologies. Existing detection methods can often only inform the existence of relics, while the risk output module of this system can mark points with relics risk by analyzing the point information of the relics abnormal point set, making the early warning results more targeted. Relics of different locations and characteristics may cause different impacts. Items left in evacuation passages may hinder the evacuation of personnel in an emergency, and items left near equipment may affect the normal operation of the equipment. By marking the risks of these points, the system allows operators to intuitively understand the potential impact of each relic, so that they can take appropriate measures to reduce unnecessary waste of resources and improve the accuracy of operation management. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a working principle diagram of the real-time visual recognition and detection system for objects left behind in a subway terminal carriage according to the present invention;

[0043] Figure 2 Flowchart of the image acquisition module;

[0044] Figure 3 Flowchart of the feature extraction module;

[0045] Figure 4Flowchart of the timing correlation module;

[0046] Figure 5 Flowchart for synchronizing and calibrating multi-view video streams. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] See also Figure 1-Figure 5 The present invention provides a real-time visual recognition and detection system for objects left in a subway terminal carriage. The specific implementation steps are as follows:

[0049] The image acquisition module obtains the surveillance video stream in the terminal carriage, collects the interior images of the carriage in the corresponding time period, associates the time nodes and marks the image sequence index to generate a carriage image sequence set; the feature extraction module extracts the image features and corresponding time indexes in the carriage image sequence set, arranges adjacent image pairs in chronological order, marks the stable and changing areas in the adjacent image pairs, and obtains a time-series associated partition annotation set; the time-series association module obtains the image frames located in the time-series stable segment of the time-series associated partition annotation set, extracts the pixel values ​​and color time series of the target area, compares the change amplitude in combination with the common features of the remains, evaluates the feature matching strength under image fluctuations, and generates a feature matching superposition analysis result; the anomaly judgment module identifies the anomaly points in the image frames whose feature matching values ​​are greater than the benchmark matching threshold and are in the time-series changing segment in the feature matching superposition analysis result to form a remains anomaly point set; the risk output module obtains all the points and corresponding point information in the remains anomaly point set, marks the points with remains risk, and generates real-time detection and risk warning results of remains in the carriage.

[0050] Example 1:

[0051] The collection of cabin image sequences consists of image eigenvalues, image time indexes, and normalized image factors. The image eigenvalues ​​quantitatively represent various visual information in the cabin interior images, covering aspects such as image texture, contours, and color distribution. The image time index accurately records the specific moment each frame was captured, facilitating subsequent time-series analysis of the image sequence. The normalized image factor is a parameter obtained by standardizing the raw image data using a specific algorithm. It eliminates differences between images due to factors such as lighting and shooting angle, making images captured at different time points comparable.

[0052] The temporal association partition annotation set specifically includes temporally stable segment annotation, temporally changing segment annotation, and the feature difference rate of adjacent image frames. Temporally stable segment annotation is used to mark segments in the image sequence where the image content remains relatively stable over a certain period of time. In these segments, the position and state of objects within the carriage change slightly. Temporally changing segment annotation corresponds to segments where the image content changes significantly. For example, when passengers get on or off the train or objects are moved, the relevant image frames will be marked as temporally changing segments. The feature difference rate of adjacent image frames is a quantitative indicator of the degree of feature change between two adjacent frames. By calculating this difference rate, we can intuitively understand the size of the difference between adjacent image frames, providing a basis for distinguishing stable segments from changing segments.

[0053] The feature matching overlay analysis results include the degree of influence of the pixel value change rate on the feature, the degree of influence of the color value change rate on the feature, and the feature matching comparison under each image fluctuation condition. The degree of influence of the pixel value change rate on the feature reflects the effect of the speed of pixel value change over time on the image features. When the pixel value changes rapidly, some features may become blurred or disappear. The degree of influence of the color value change rate on the feature reflects the impact of the color change speed on the image features. Different color change rates may make previously obvious features less prominent. The feature matching comparison under each image fluctuation condition compares the image features under different fluctuation conditions with the common features of the relics, thereby analyzing the feature matching under various fluctuation conditions.

[0054] The set of relic anomaly points includes the anomaly's time location, the color variation characteristics of the anomaly's pixels, and the ratio of the anomaly's characteristics to the baseline. The anomaly's time location accurately records the time corresponding to the image frame identified as the anomaly, helping personnel trace the specific moment of the anomaly. The color variation characteristics of the anomaly describe the color variation characteristics of the pixel at the anomaly over a certain period of time. By analyzing this characteristic, we can understand the color variation pattern of the anomaly. The anomaly's characteristics to the baseline fluctuation ratio is the ratio of the anomaly's characteristic variation to the preset baseline fluctuation amplitude, which reflects the degree of deviation of the anomaly's characteristic change.

[0055] The real-time detection and risk warning results for leftovers in train carriages include a list of detected anomaly points and a label based on the three combined indicators of the anomaly points. The list details the specific locations of all detected anomalies within the carriage, making it easier for personnel to quickly locate them. The label based on the three combined indicators of the anomaly points is generated by comprehensively considering multiple characteristic indicators of the anomaly points, allowing for a more comprehensive assessment of the nature of the anomaly points.

[0056] The video stream acquisition submodule within the image acquisition module captures the surveillance video stream data from the terminal carriage. During this acquisition process, it simultaneously captures the time nodes in the video stream and the corresponding carriage interior images. These acquisition results are recorded as two basic factors: the time factor and the image factor, thereby forming the basic carriage image data set. The time factor is directly related to the time information of image acquisition, while the image factor contains the raw data information of the carriage interior images.

[0057] The image preprocessing submodule performs normalization on the time and image factors in the basic vehicle image dataset. Normalization of the time factor converts time information of varying formats and precision into a unified time representation to facilitate time series analysis. Normalization of the image factor adjusts the image size, brightness, contrast, and other factors to maintain a consistent standard. The normalized results are mapped to time nodes, and the average of the normalized time and image values ​​is calculated as the image feature value, ultimately generating a vehicle image sequence.

[0058] The video stream acquisition submodule also captures video stream data from multiple surveillance cameras within the terminal carriage. Due to the varying installation positions and angles of the cameras, the captured video stream data may exhibit temporal discrepancies. Therefore, the video streams require synchronization and calibration to ensure accurate correspondence between time nodes and image frames. Simultaneously, images of the carriage interior are collected from various angles and recorded as multi-angle image factors. By integrating these multi-angle image factors, a basic multi-view carriage image data set is generated, providing a more comprehensive view of the carriage interior and reducing potential blind spots or misjudgments caused by single-angle shooting.

[0059] Example 2:

[0060] The feature extraction module includes a feature extraction submodule, a feature comparison submodule and a trend classification submodule. Each submodule works together to complete the feature extraction and partition labeling of the carriage image sequence.

[0061] The feature extraction submodule first obtains image feature values ​​and corresponding time data from the carriage image sequence. The image feature values ​​encompass various visual features relevant to artifact recognition, such as object outlines, texture distribution, and color mean. The time data is accurate to the acquisition time of each frame, including specific information such as year, month, day, hour, minute, and second. Based on this time information, the feature extraction submodule identifies the temporal relationship of all image frames, determining which frames were acquired first and which were acquired later, thereby forming a chronologically ordered image sequence framework. Subsequently, the image time index set is called, which records the sequence number and corresponding timestamp of each frame within the entire video stream, facilitating rapid image location and association. Time is then measured and sorted based on a time threshold. The time threshold is a set time interval, such as 0.5 seconds. Two frames are considered adjacent if the acquisition time interval is less than or equal to this threshold. This generates a time-ordered sequence of adjacent image frames, clearly displaying the order of adjacent frames and their corresponding time information.

[0062] The feature comparison submodule operates based on a temporally ordered sequence of adjacent image frames. For the i and jth adjacent image frames in the sequence, grayscale conversion is first performed. Grayscale conversion converts a color image into a grayscale image. The grayscale value of each pixel is calculated by calculating the weighted average of the red, green, and blue color components. This simplifies the image data and reduces the computational complexity of subsequent processing. Next, edge contour features are extracted from the grayscaled image. Edge contours are the interface between an object and its background in an image, typically appearing as areas with dramatic changes in grayscale value. Specific edge detection algorithms can identify these contours, which can include straight lines, curves, and combinations thereof. The similarity of the edge contours of the two adjacent image frames is calculated, serving as the basis for evaluating the feature difference rate. The calculation of contour similarity involves comparing contour shape, length, and positional distribution. When the similarity is high, the feature difference rate is low, indicating minimal change between the two image frames. When the similarity is low, the feature difference rate is high, indicating significant differences between the two image frames. The feature difference rates of all adjacent image frame pairs are integrated to generate a feature difference rate sequence including contour similarity, which fully records the degree of feature difference between each pair of adjacent image frames.

[0063] The trend classification submodule further extracts pixel and color change trends between adjacent image frames based on the feature difference rate sequence. Pixel change trends refer to the overall direction of change in pixel values ​​between two adjacent frames, for example, whether pixel values ​​generally increase, decrease, or show no significant overall change. For color images, color change trends analyze the direction of change in parameters such as hue, saturation, and brightness of the primary colors between two adjacent frames. For example, whether the overall hue leans toward warm or cool, or whether the saturation increases or decreases. Each pair of image frames is classified and labeled based on whether these two trend change directions are consistent. When the pixel and color change trends are in the same direction, i.e., both increase or decrease simultaneously, the trend is considered consistent. When the two change in opposite directions, such as when pixel values ​​increase while color saturation decreases, the trend is considered conflicting. Segments with consistent and conflicting trends are recorded and grouped separately. Segments with consistent trends may indicate stable changes in the vehicle cabin environment, while segments with conflicting trends may indicate the presence of an unusual object or a sudden change in the environment. Through these operations, we finally obtain a time-series correlation partition annotation set, which clearly divides the time-series stable segments and the time-series changing segments, providing a structured image partition basis for subsequent relic identification.

[0064] Throughout the entire process, the processing results of each submodule are interconnected and mutually supportive. The time-ordered sequence of adjacent image frames generated by the feature extraction submodule provides clear processing targets for the feature comparison submodule; the feature difference rate sequence calculated by the feature comparison submodule provides quantitative data for trend analysis in the trend classification submodule; and the classification annotation results of the trend classification submodule serve as the core content of the temporal correlation partition annotation set. This layered processing approach accurately captures the changing patterns of image features in a continuous sequence of carriage images, laying the foundation for the subsequent identification of abnormal points of leftover objects. Furthermore, the focus on edge contour features enables the stability of contours to assist in determining differences in image features, even in complex environments such as changing lighting, thereby improving the reliability of feature difference rate calculations. Trend classification partitions the image sequence based on the overall direction of change, helping to distinguish normal environmental changes from abnormal changes that may be caused by leftover objects, making the entire feature extraction process more targeted and effective.

[0065] Example 3:

[0066] The time series association module includes a sequence extraction submodule and a change rate calculation submodule. It generates feature matching and superposition analysis results by analyzing image frames in the time series stable segment.

[0067] The sequence extraction submodule first receives a set of time-correlated partition annotations, which clearly mark temporally stable and temporally changing segments. A temporally stable segment refers to a period in which the overall environment and distribution of major objects within the train compartment do not change significantly across multiple consecutive image frames. For example, this period occurs after the train stops at the last station, all passengers have disembarked, and no new passengers enter. The sequence extraction submodule selects these segments marked as temporally stable and, for each image frame within these segments, detects the pixel values ​​and relative color value data of the target area within the acquisition time period. The target area refers to areas within the train compartment where objects may be left behind, such as seat surfaces, floors, and luggage racks, and is defined by a pre-defined coordinate range. The pixel value refers to the brightness of each pixel within the target area, typically expressed as an integer between 0 and 255. The relative color value is the value obtained by converting the color information of the target area into a specific color space (such as HSV space), consisting of three components: hue, saturation, and lightness. Arrange these pixel values ​​in chronological order to form a pixel time series. Similarly, arrange the relative color values ​​in chronological order to form a color time series. The two together constitute an image change time series set.

[0068] The rate of change calculation submodule operates based on the image change time series set. For the pixel time series, the rate of change of the pixel values ​​between consecutive time nodes in each image frame time series is calculated. Consecutive time nodes refer to the acquisition moments of two adjacent image frames, such as time t1 and time t2, where t2 is later than t1. The pixel value change rate is calculated by subtracting the average pixel value of the target area at time t1 from the average pixel value of the target area at time t2, and then dividing it by the time interval between the two moments. The formula is as follows: in, Indicates the rate of change of pixel value, Indicates the The average pixel value of the target area at the moment, Indicates the The average pixel value of the target area at the moment, and Respectively represent the collection moments of two consecutive time nodes.

[0069] Similarly, for the color time series, the rate of change of the color value between consecutive time nodes is calculated. Taking the hue component in the HSV space as an example, the change rate formula is: in, Indicates the rate of change of hue value, Indicates the The average hue value of the target area at the moment, Indicates the The average hue value of the target area at the moment, and The meaning of is the same as above. For the saturation and lightness components, the change rate is calculated in a similar way, which is recorded as (rate of change of saturation) and (Rate of change of brightness).

[0070] The pixel value change rate and color value change rate (including 、 、 ) and conduct side-by-side comparisons under the same characteristic conditions. The same characteristic conditions refer to changes within the same target area and the same time period. By jointly analyzing these two types of rate indicators, we can identify their numerical relationship to the magnitude of feature changes. For example, when the pixel value change rate is positive and large, and the hue value change rate is also positive and large, it indicates that the target area may have experienced a simultaneous increase in brightness and hue. When the pixel value change rate is close to zero, and the saturation change rate is negative and small, it indicates that the brightness of the target area remains essentially unchanged, but the color vividness has slightly decreased.

[0071] The influence value sequence for each image frame is integrated. The influence value sequence is formed by quantifying the degree of influence of the pixel value change rate and the rate of change of various color values ​​on the features. The degree of influence of the pixel value change rate on the features is determined by analyzing the effect of pixel value changes on the target area's contour clarity and texture characteristics. When the pixel value change rate is too large, it may lead to blurred contours, thus affecting feature recognition. The degree of influence of the color value change rate on the features is analyzed by analyzing the effect of color changes on the distinction between the target area and the background. Excessive color change rates can cause the original color features to become unstable. Feature matching and comparison under each image fluctuation condition is to compare the target area features under the current fluctuation conditions (such as a specific combination of pixel and color change rates) with the features in the common feature library of relics, and count the number of matched features and the degree of matching.

[0072] Through the above process, a feature matching overlay analysis result is established. This result integrates the impact of pixel and color change rates on features, as well as feature matching under different fluctuation conditions. It can clearly demonstrate how the characteristics of the target area change over time during the stable time series segment, and the degree of match between these changes and the characteristics of the remnants. This analysis method can effectively filter out feature fluctuations caused by normal factors such as slow changes in light and slight camera jitter, focusing on feature changes that may be caused by the presence of remnants, providing detailed feature change data for subsequent anomaly determination. At the same time, the quantitative analysis of the change rate provides a clear numerical representation of the speed and magnitude of feature changes, facilitating precise comparison and judgment, and ensuring the accuracy of subsequent anomaly identification.

[0073] Example 4:

[0074] The anomaly determination module consists of a record extraction submodule, a parameter change ratio calculation submodule and an anomaly identification submodule. It generates a set of abnormal points of relics by gradually processing the results of feature matching and superposition analysis.

[0075] After receiving the results of the feature matching overlay analysis, the record extraction submodule first selects two types of image frames: those with feature matching values ​​greater than a baseline matching threshold, indicating areas with features similar to those commonly found in objects; and those within a time-varying period, indicating significant changes in the environment or object distribution within the vehicle. These selected image frames are then continuously captured and recorded in chronological order, specifically including pixel and color value data for the target area corresponding to each time point. The target area is pre-defined based on the vehicle's structure, encompassing areas prone to object retention, such as seats, floors, and armrests. Pixel values ​​reflect the brightness of the area, while color values ​​include details such as hue and saturation. This data is then organized chronologically to form a continuous acquisition record set, which fully captures all relevant image frame data from the beginning to the end of the changing period when the feature matching criteria are met.

[0076] The parameter change ratio calculation submodule calls the continuous acquisition record set and refines the image frame data. The pixel values ​​and color values ​​of two consecutive time nodes are extracted from the record set, such as Moment and moment, in which for When calculating the pixel change ratio, take the absolute difference between the pixel values ​​at these two moments and then add the difference to the first The pixel value at the moment is divided by the pixel value at the moment, and the result is the pixel change ratio between the two moments. Similarly, when calculating the color change ratio, use the Moment and The absolute difference of the color value at the moment is divided by the The color value at each moment is used to obtain the color change ratio. The pixel change ratios for all consecutive time node pairs are arranged sequentially to form a pixel change ratio sequence. Similarly, the color change ratios are arranged sequentially to form a color change ratio sequence. These two sequences, along with the corresponding time node information, form an image fluctuation dataset, which intuitively presents the magnitude of pixel and color changes in the target area over continuous time.

[0077] The anomaly recognition submodule identifies outliers based on the image fluctuation dataset. Pixel distribution data and color distribution data for corresponding time periods are extracted from the dataset. The pixel distribution data indicates the frequency of pixel values ​​within the target area, while the color distribution data reflects the distribution of each color component. A fluctuation recognition threshold is set based on the range of pixel and color variations under normal environmental changes in historical data. First, the pixel and color change ratios at consecutive time points are determined to exceed the set fluctuation recognition threshold. If both exceed the threshold, the magnitude of pixel and color changes in the target area exceeds the normal range. Next, the pixel and color distributions are determined to deviate from the baseline distribution range. The baseline distribution range is calculated based on a large number of carriage images without debris. If both deviate, the overall visual characteristics of the target area are significantly different from the normal state. When a time point meets both of these conditions, it is marked as an outlier. All marked outliers are sorted chronologically to form a debris anomaly point set. This set contains each outlier's corresponding time position, pixel color change amplitude characteristics, and comparison information with the baseline fluctuation.

[0078] Throughout the entire process, the processing logic of each submodule is tightly integrated. The record extraction submodule filters and focuses on areas of change that may contain artifacts, narrowing the scope for subsequent analysis. The parameter change ratio calculation submodule quantifies the magnitude of change, converting visual image changes into comparable numerical values, making the degree of change more intuitive. The anomaly identification submodule uses dual criteria to ensure that the marked anomalies both exceed the normal range in magnitude and deviate from the baseline in overall distribution, thereby reducing misidentification caused by fluctuations in a single factor. For example, when a passenger leaves a backpack in a subway terminal carriage, the pixel values ​​in the area where the backpack is located differ from those in the surrounding environment in consecutive image frames. From the time the passenger leaves until the backpack comes to rest, the pixel change ratio and color change ratio in this area may suddenly increase and exceed the fluctuation detection threshold. At the same time, the pixel and color distribution of this area may also deviate from the baseline range when no artifacts are left behind. At this time, the corresponding time point is marked as an anomaly and included in the artifact anomaly point set. This layered processing approach accurately captures anomalous areas that may be artifacts from continuous image changes.

[0079] Example 5:

[0080] The risk output module includes an indicator joint judgment submodule and an abnormal output sorting submodule. By analyzing and processing the abnormal point set of the leftovers, it generates real-time detection and risk warning results of leftovers in the carriage.

[0081] The indicator joint judgment submodule obtains all points in the abnormal point set of the remains and the corresponding time index and location information. The time index of each point is accurate to the acquisition time of the image frame, including the specific hour, minute and second, and the location information is marked by the preset coordinate system in the carriage, such as a row of seats in a certain carriage, a certain ground area, etc. For each point, the joint risk judgment value of the kth image frame is calculated. This value comprehensively considers multiple parameters such as the pixel change ratio, color change ratio, feature matching value, etc. of the abnormal point, and obtains a comprehensive value through the weighted combination of these parameters. The joint risk judgment values ​​of all points are arranged in chronological order to form a joint risk judgment value sequence, which reflects the risk level distribution of abnormal points at different time points.

[0082] The abnormal output sorting submodule further screens points with legacy risks based on the joint risk judgment value sequence. The feature matching risk threshold and feature reference value are set. The feature matching risk threshold is set according to the feature difference between common relics and non-relic objects, and the feature reference value is determined based on the image feature value of the normal area in the car when there are no relics. The screening conditions include: the feature matching value is greater than the feature matching risk threshold, indicating that the feature of the point has a high degree of match with the feature of the relics; the image feature value is lower than the feature reference value, indicating that the image feature of the point is significantly different from the normal area; the time series association label is a changing segment, which means that the car environment has changed during the time period of the point, which is consistent with the scenario where relics may appear. Points that meet all three conditions are marked as detected anomalies and have legacy risks.

[0083] For the selected points, extract their corresponding time index, location identifier, and partition. The time index is used to trace the specific moment when the anomaly occurred. The location identifier specifies the specific location of the point in the car through coordinates or area names. The partition is divided according to the car structure, such as seating area, aisle area, door area, etc. This information is partitioned and output in a structured format. The structured format includes a table or list, which contains fields such as point number, time, location, partition, risk level, etc. When outputting the partitions, the points are classified according to the partitions they belong to, and the points in the same partition are displayed together to facilitate quick positioning and processing by staff.

[0084] During the output process, the information of each point is checked to ensure the accuracy of the time index and location identification, and to avoid misjudgment due to data errors. For multiple abnormal points that appear at the same location at different times, they are arranged in chronological order to show their risk change trends. The generated real-time detection and risk warning results of leftovers in the carriage also include a joint judgment label of the three indicators of the abnormal point. This label combines the judgment results of the feature matching value, image feature value and time series association label, and is expressed in concise symbols or text form, such as "high match-low feature-change area", which intuitively reflects the key features of each abnormal point.

[0085] Throughout the risk output module's processing, the indicator joint determination submodule and the abnormal output collation submodule work together to ensure accurate and easy-to-understand detection and warning results through multi-condition screening and structured output. Based on the output results, staff can quickly identify the possible locations and corresponding time information of leftovers within the carriage, allowing for timely processing and reducing the risk of missed items. Furthermore, the partitioned output method improves processing efficiency, enabling staff to conduct a zone-by-zone investigation to ensure coverage of all areas where leftovers may be present.

[0086] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0087] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A real-time visual recognition and detection system for objects left behind in a subway terminal carriage, characterized in that: The system comprises: The image acquisition module obtains the surveillance video stream inside the terminal carriage, collects the carriage interior images of the corresponding time period, associates the time nodes and marks the image sequence index, and generates a carriage image sequence set; The feature extraction module extracts image features and corresponding time indexes from the carriage image sequence set, arranges adjacent image pairs in chronological order, marks stable and changing regions in the adjacent image pairs, and obtains a temporal correlation partition annotation set; The temporal association module obtains the image frames located in the temporally stable segment of the temporal association partition annotation set, extracts the pixel values ​​and color time series of the target area, compares the change amplitude with the common features of the remains, evaluates the feature matching strength under image fluctuations, and generates feature matching superposition analysis results; The anomaly determination module identifies anomalies in the image frame whose feature matching value is greater than a reference matching threshold and is in a time sequence change section in the feature matching superposition analysis result, and forms a set of remnant anomaly points; The risk output module obtains all points in the abnormal point set of the leftover objects and the corresponding point information, marks the points with leftover risks, and generates real-time detection and risk warning results of leftover objects in the carriage; The image acquisition module includes: The video stream acquisition submodule acquires the surveillance video stream data in the terminal carriage, collects the time nodes and corresponding carriage interior images in the video stream, and records the acquisition results as two basic factors, time factor and image factor, to obtain the carriage image basic data set; The image preprocessing submodule performs normalization processing on the time factor and image factor data in the carriage image basic data set, establishes a correspondence between the normalized results and the time nodes, calculates the average of the normalized time value and the normalized image value as the image feature value, and generates a carriage image sequence set; The feature extraction module includes: The feature extraction submodule obtains the image feature values ​​and corresponding time data in the carriage image sequence set, identifies the positional relationship of all image frames in the time sequence based on the time information, calls the image time index set, and uses the adjacent time threshold as a reference to perform time measurement and sort the image frames in the time sequence to generate a time sorted sequence of adjacent image frames; The feature comparison submodule calculates the feature difference rate between the i-th and j-th adjacent image frames based on the time-ordered sequence of the adjacent image frames, and integrates them to generate a feature difference rate sequence; The trend classification submodule extracts pixel change trends and color change trends between adjacent image frames based on the feature difference rate sequence, classifies and labels each pair of image frames based on whether the two trend change directions are consistent, and records and groups the segments with consistent and conflicting trends respectively to obtain a time-series associated partition annotation set; The timing association module includes: The sequence extraction submodule selects the segments marked as time-stable according to the time-series associated partition annotation set, detects the pixel values ​​and relative color value data of the target area within each image frame acquisition time period, arranges them in chronological order to form pixel time series and color time series, and generates an image change time series set; The change rate calculation submodule calculates the pixel value change rate and color value change rate between consecutive time nodes in each image frame time series based on the image change time series set, compares the pixel value change rate and color value change rate in parallel under the same feature conditions, identifies the numerical relationship of the change amplitude under the feature by jointly analyzing the two types of rate indicators, integrates the influence value sequence of each image frame, and establishes the feature matching superposition analysis result.

2. The real-time visual recognition and detection system for objects left behind in a subway terminal carriage according to claim 1 is characterized in that: The carriage image sequence set includes image feature values, image time indexes and normalized image factors. The time-series associated partition annotation set specifically includes time-series stable segment annotations, time-series changing segment annotations and feature difference rates of adjacent image frames. The feature matching superposition analysis results include the degree of influence of the pixel value change rate on the feature, the degree of influence of the color value change rate on the feature and the feature matching comparison under each image fluctuation condition. The residue anomaly point set includes the time position of the anomaly point, the color variation feature of the anomaly point pixel and the ratio of the anomaly point feature to the benchmark fluctuation. The real-time detection and risk warning results of residues in the carriage include a list of detected anomaly points and a joint judgment label of the three indicators of the anomaly point.

3. The real-time visual recognition and detection system for objects left behind in a subway terminal carriage according to claim 1 is characterized in that: The abnormality determination module includes: The record extraction submodule selects image frames whose feature matching values ​​are greater than the reference matching threshold and image frames in the time-series change section based on the feature matching overlay analysis results, extracts the continuous acquisition records of the image frames in chronological order, collects the pixel value and color value data of the target area corresponding to each time node, and generates a continuous acquisition record set; The parameter change ratio calculation submodule calls the continuous acquisition record set, extracts the pixel values ​​and color values ​​of the image frame at two consecutive time nodes, calculates the pixel change ratio and color change ratio respectively, integrates them into a pixel change ratio sequence and a color change ratio sequence, and establishes an image fluctuation change data set; The anomaly recognition submodule extracts the pixel distribution data and color distribution data of the corresponding time period based on the image fluctuation change data set, determines whether the pixel change ratio and color change ratio both exceed the set fluctuation recognition threshold, and determines whether the pixel distribution and color distribution deviate from the baseline distribution range at the same time. The time nodes that meet the conditions are marked as anomalies to generate a set of relic anomaly points.

4. The real-time visual recognition and detection system for objects left in a subway terminal carriage according to claim 3 is characterized in that: The risk output module includes: The indicator joint judgment submodule obtains all points in the residue abnormal point set and the corresponding time index and position information, calculates the joint risk judgment value of the k-th image frame, and establishes a joint risk judgment value sequence; Based on the joint risk judgment value sequence, the abnormal output sorting submodule screens points whose feature matching values ​​are greater than the feature matching risk threshold, image feature values ​​are lower than the feature reference value, and whose time series association labels are change segments, extracts the corresponding point time index, location identifier and partition to which it belongs, marks them as detected abnormalities with residual risks, and outputs the points that meet the joint conditions in a structured format, generating real-time detection and risk warning results for residual objects in the carriage.

5. The real-time visual recognition and detection system for objects left behind in a subway terminal carriage according to claim 1 is characterized in that: The video stream acquisition submodule includes acquiring video stream data from multiple surveillance cameras in the terminal carriage, synchronously calibrating the video stream to ensure the accurate correspondence between time nodes and image frames, collecting carriage interior images from different angles and recording them as multi-angle image factors, and obtaining a multi-view carriage image basic data set.

6. The real-time visual recognition and detection system for objects left behind in a subway terminal carriage according to claim 1 is characterized in that: The feature comparison submodule includes grayscale processing of adjacent image frames, extracting image edge contour features, calculating contour similarity as a basis for evaluating feature difference rate, and integrating to generate a feature difference rate sequence including contour similarity.

7. The real-time visual recognition and detection system for objects left in a subway terminal carriage according to claim 3 is characterized in that: The parameter change ratio calculation submodule includes calculating the ratio of the absolute difference between the pixel values ​​of two consecutive time nodes and the pixel value of the previous time node as the pixel change ratio, and calculating the ratio of the absolute difference between the color values ​​of two consecutive time nodes and the color value of the previous time node as the color change ratio, and integrating them into a pixel change ratio sequence and a color change ratio sequence containing the absolute difference ratio.

Citation Information

Patent Citations

  • Detection method for remnants in complex environment

    CN104156942A

  • A remaining object detection method and device

    CN109948455A