A bus passenger attribute selective recognition system, method and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]这种“不可见属性被误输出为具体类别”的问题,会导致基于属性的乘客检索系统产生系统性假阳性检索结果
[0072]第一,本发明通过引入属性级可观测性预测分支,显式地预测每个属性在当前图像中的可见状态(可见、部分可见、不可见),将“属性不可见”作为一个独立的、可量化的因素纳入决策过程,与现有方法中将不可见属性简单忽略或统一归为不确定性的做法相比,能够更准确地反映属性识别的实际可靠程度。
Smart Images

Figure CN122551331A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent public transportation and image recognition, and in particular to a selective output recognition system, method and medium for bus passenger attributes. Background Technology
[0002] Bus surveillance video is a crucial data source for urban public transportation safety management, structured passenger retrieval, passenger flow analysis, and incident tracing. Passenger attribute recognition aims to extract semantic attributes such as gender, age group, clothing type, clothing color, mask, glasses, hat, backpack, shoes, and carried items from surveillance images.
[0003] Existing feature attribute recognition methods typically model the task as a forced output problem of "outputting a specific category for each feature attribute." However, factors such as seats, handrails, other passengers obstructing the view in a bus, camera perspective, low resolution, and motion blur can make certain local attributes invisible or only partially visible in the current frame. For example, the shoe area is often obscured by seats, and the face area may be obscured by masks, other passengers, or the perspective. In such cases, if the system still outputs a specific category, the output result is often generated by prior bias or model guessing, rather than supported by the current image evidence.
[0004] This problem of "invisible attributes being mistakenly output as specific categories" can lead to systematic false positive search results in attribute-based passenger retrieval systems. For example, when searching for "passengers wearing blue masks," the system may incorrectly return a passenger whose face is completely obscured, making it impossible to determine the color or condition of the mask. In security audits and incident retrospectives, this type of error wastes human review resources and may cause the true target to be missed.
[0005] Therefore, there is an urgent need for a method that enables the system not only to predict attribute categories, but also to determine whether each feature attribute is observable in the current frame, and to reject the output of a specific category when the risk is high, thereby reducing the probability of invisible attributes being mistakenly output as specific categories and improving the reliability of bus passenger attribute recognition results. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a selective identification system and method for passenger attributes in public transportation. Instead of simply modeling passenger attribute identification as a forced classification problem of "input image to specific attribute category", it further introduces attribute-level observability, uncertainty and reliability decision-making, so that the system can determine whether each attribute has sufficient visual evidence in the current frame and output an unknown state when the risk is high.
[0007] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows:
[0008] A selective identification system for bus passenger attributes includes an image acquisition and preprocessing module, an attribute identification and observability output module, an uncertainty estimation unit, a risk calculation and calibration unit, and a selective output unit.
[0009] The image acquisition and preprocessing module obtains a single passenger image;
[0010] This is used to acquire bus surveillance video images, determine passenger target areas from the bus surveillance video images, and generate passenger images to be identified based on the passenger target areas.
[0011] The attribute recognition and observability output module outputs the attribute prediction category and the attribute-level observable state probability vector, respectively.
[0012] This is used to receive the image of the passenger to be identified, and to output the attribute category probability distribution and predicted category, as well as the attribute-level observable state probability vector, for each passenger attribute; wherein, for the i-th passenger attribute, the attribute category probability distribution is expressed as... The attribute-level observable state probability vector includes the visibility probability. Partially visible probability and invisibility probability "Visible" indicates that the visual area corresponding to the attribute is clearly visible; "partially visible" indicates that the area corresponding to the attribute is partially visible or partially occluded; "invisible" indicates that the area corresponding to the attribute is invisible or lacks reliable visual evidence.
[0013] The uncertainty estimation unit calculates attribute-level uncertainty;
[0014] Based on the attribute category probability distribution of the i-th passenger attribute Computational property level uncertainty ;
[0015] The probability of feature attributes being invisible and attribute level uncertainty Merge into attribute-level risk scores :
[0016] ;
[0017] in, and These represent the risk weights for the probability of an attribute being invisible and the uncertainty, respectively.
[0018] The risk calculation and calibration unit determines the attribute-level risk score and calibration threshold of the characteristic attribute; based on the attribute-level uncertainty... and invisibility probability Determine the attribute level risk score s for the i-th passenger attribute. iAnd based on the validation set, learn the attribute-level risk calibration threshold for each feature attribute. ;
[0019] The selective output unit is used to output a specific attribute category or an unknown state;
[0020] If the risk score of the i-th feature attribute satisfies If the attribute classification branch outputs a specific attribute category, then the attribute is considered a reliable output, and the predicted category is output. ;
[0021] like If the attribute is determined to lack reliable visual evidence, then the attribute is determined to be an unreliable output, and the output state is unknown.
[0022] In the aforementioned selective identification system for bus passenger attributes, preferably, the risk calculation and calibration unit is used to determine the corresponding attribute-level risk calibration threshold for the i-th passenger attribute based on the attribute-level risk score and the actual attribute-level observable state label in the validation set. :
[0023] The risk calculation and calibration unit calculates the false output rate of the specific category of the invisible attribute for the i-th passenger attribute based on the number of samples in the verification set whose attribute-level risk scores are not greater than the candidate threshold τ.
[0024] ;
[0025] The risk calculation and calibration unit calculates the effective attribute coverage of the i-th passenger attribute based on the number of samples in the verification set whose attribute-level risk scores are not greater than the candidate threshold τ, both in the truly visible samples and the truly partially visible samples.
[0026] = ;
[0027] The risk calculation and calibration unit ensures that the false output rate of the specific category of the invisible attribute is not greater than the preset upper limit of the false output rate of the specific category of the invisible attribute. Under the condition of, from the candidate threshold set of the i-th passenger attribute Select the candidate threshold that maximizes the coverage of the effective attributes, and use it as the attribute-level risk calibration threshold for the i-th passenger attribute:
[0028] ;
[0029] in, Indicates the number of samples in the validation set. This represents the true observable state label of the i-th passenger attribute in the n-th validation sample. This indicates the corresponding attribute-level risk score. Represents the set of candidate thresholds. This represents the upper limit of the error output rate for a specific category of the preset invisible attribute, and 1(⋅) is the indicator function.
[0030] In the aforementioned selective identification system for bus passenger attributes, preferably, the uncertainty estimation unit is used to estimate the probability distribution of the attribute category of the i-th passenger attribute. Computational property level uncertainty
[0031] Wherein, the attribute category probability distribution of the i-th passenger attribute Represented as:
[0032] ;
[0033] Classification confidence of the i-th passenger attribute for:
[0034] ;
[0035] The attribute level uncertainty Calculate according to the following formula:
[0036] ;
[0037] in, This represents the number of candidate attribute categories for the i-th passenger attribute. This represents the probability that the i-th passenger attribute belongs to the k-th candidate attribute category. The larger the value, the more uncertain the model's prediction of the category of the i-th passenger attribute.
[0038] In the aforementioned selective identification system for bus passenger attributes, preferably, the image acquisition and preprocessing module includes an image acquisition unit and a passenger target area generation unit;
[0039] The image acquisition unit is used to acquire monitoring images of the bus compartment;
[0040] The passenger target area generation unit is used to determine at least one passenger target area from the bus carriage monitoring image, and generate a corresponding passenger image to be identified based on the passenger target area.
[0041] In the aforementioned selective identification system for bus passenger attributes, preferably, the attribute identification and observability output module includes a feature extraction unit, a coarse-grained region pooling unit, an attribute classification branch, and an attribute-level observability prediction branch.
[0042] The feature extraction unit is used to receive the passenger image to be identified and extract the shared feature map F from the passenger image to be identified;
[0043] The coarse-grained region pooling unit is used to perform coarse-grained region pooling on the shared feature map F to obtain the region features corresponding to the passenger's body region;
[0044] The attribute classification branch is used to output the attribute category probability distribution and predicted category for each passenger attribute based on the shared feature map F and the region features;
[0045] The attribute-level observability prediction branch, based on shared features and regional features, is used to output attribute-level observable state probability vectors for each passenger attribute based on the shared feature map F and the regional features; wherein, for the i-th passenger attribute, the attribute-level observable state probability vector includes the visibility probability. Partially visible probability and invisibility probability .
[0046] In the aforementioned selective identification system for bus passenger attributes, preferably, the attribute-level observability prediction branch is used to predict the attribute-level observability state of each passenger attribute in the current image of the passenger to be identified.
[0047] For the i-th attribute, the attribute-level observability prediction branch outputs an observable state prediction score vector:
[0048] ;
[0049] The Softmax function is then used to convert the observable state prediction score vector into an attribute-level observable state probability vector.
[0050] ;
[0051] in,( ), ( )and( ) represent the probabilities of the i-th passenger attribute being in a visible state, a partially visible state, and an invisible state, respectively; the visible state means that the visual area corresponding to the passenger attribute is clearly visible, the partially visible state means that the visual area corresponding to the passenger attribute is partially visible or partially occluded, and the invisible state means that the visual area corresponding to the passenger attribute is invisible or lacks reliable visual evidence;
[0052] The invisibility probability of the i-th passenger attribute is the probability component of the corresponding invisible state in the observable state probability vector of the attribute level:
[0053] ;
[0054] in, This represents the observable state of the i-th passenger attribute. This indicates the image of the passenger to be identified.
[0055] A method for selectively identifying passenger attributes on public transport includes the following steps:
[0056] 1) Acquire bus cabin surveillance images, determine passenger target areas from the bus cabin surveillance images, and generate passenger images to be identified based on the passenger target areas;
[0057] 2) Extract features from the image of the passenger to be identified to obtain a shared feature map F; perform coarse-grained region pooling on the shared feature map F to obtain region features related to the passenger's body region; output the predicted category and confidence score of each passenger attribute using the attribute classification branch; output the visibility probability, partial visibility probability, and invisibility probability of each passenger attribute using the observability prediction branch, where the invisibility probability is denoted as... ;
[0058] 3) Calculate the attribute level uncertainty based on the attribute category probability distribution of the passenger attributes. The probability of the attribute being invisible. and attribute level uncertainty Merge into attribute-level risk scores :
[0059] ;
[0060] in, and These represent the risk weights for the probability of an attribute being invisible and the uncertainty, respectively.
[0061] 4) Determine the attribute-level risk calibration threshold for each passenger attribute based on the validation set. ;
[0062] 5) If the risk score of the i-th passenger attribute satisfies If the attribute classification branch outputs a specific attribute category, then the attribute is considered a reliable output, and the predicted category is output. ;
[0063] like If the attribute is determined to lack reliable visual evidence, then the attribute is determined to be an unreliable output, and the output state is unknown.
[0064] In the above-mentioned selective identification method for bus passenger attributes, preferably, the attribute identification and observability output module includes a feature extraction unit, a coarse-grained region pooling unit, an attribute classification branch, and an attribute-level observability prediction branch;
[0065] Step 2) includes the following steps:
[0066] 2.1) The feature extraction unit takes the passenger image to be identified as input and obtains a shared feature map F;
[0067] 2.2) The coarse-grained region pooling unit performs coarse-grained region pooling on the shared feature map F to obtain region features related to the passenger's body region;
[0068] 2.3) The attribute classification branch, based on the shared feature F and the regional feature, outputs the category probability distribution and predicted category for each passenger attribute;
[0069] 2.4) The attribute-level observability prediction branch outputs the visibility probability, partial visibility probability, and invisibility probability of each passenger attribute based on the shared feature F and the regional feature.
[0070] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for selectively identifying public transport passenger attributes.
[0071] Compared with the prior art, the present invention has the following advantages:
[0072] First, by introducing an attribute-level observability prediction branch, this invention explicitly predicts the visibility state (visible, partially visible, invisible) of each attribute in the current image, incorporating "attribute invisibility" as an independent and quantifiable factor into the decision-making process. Compared with existing methods that simply ignore or uniformly classify invisible attributes as uncertainty, this invention can more accurately reflect the actual reliability of attribute recognition.
[0073] Second, this invention weights and fuses attribute-level observability (invisibility probability) and attribute-level uncertainty (complement of classification confidence) to construct an attribute-level risk score. This risk score simultaneously considers two sources of unreliability: the lack of visual evidence at the data level and insufficient discriminative ability at the model level, enabling a more comprehensive assessment of the risk level of each attribute identification result.
[0074] Third, this invention achieves attribute-level adaptive selective output by independently determining the optimal risk calibration threshold for each attribute through a risk calibration mechanism based on the validation set. The system can maximize the coverage of effective attributes while keeping risks under control, avoiding the problem of large performance differences among different attributes inherent in traditional fixed-threshold methods.
[0075] Fourth, in the typical scenario of bus cabin monitoring, this invention addresses practical problems such as severe obstruction, unobservable local attributes, and low resolution by employing a selective output mechanism. This mechanism enables the system to proactively output an "unknown state" rather than forcibly outputting a specific category when reliable visual evidence is lacking. This improves the credibility of passenger attribute identification results and provides more reliable data support for downstream tasks such as attribute-based passenger retrieval, security review, and event backtracking. Attached Figure Description
[0076] Figure 1 This is a flowchart of the selective identification system for bus passenger attributes in this invention.
[0077] Figure 2 This is a module architecture diagram of the selective identification system for bus passenger attributes in this invention.
[0078] Figure 3 This is a flowchart of the attribute-level observability prediction branch processing in this invention.
[0079] Figure 4 This is a comparison chart showing the specific category error rates of different methods on invisible attributes. Detailed Implementation
[0080] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the following embodiments are for illustration and explanation only and are not intended to limit the scope of protection of the present invention. Where there is no conflict, the following embodiments and their technical features can be combined with each other.
[0081] It should be noted that, in this invention, "module" and "unit" both refer to functional modules, which can be implemented by software programs, hardware circuits, or a combination of software and hardware. The functions of multiple modules or units can be implemented by the same processor executing a computer program stored in memory. Unless otherwise defined, all technical terms used below have the same meaning as commonly understood by those skilled in the art.
[0082] This invention proposes a selective passenger attribute recognition system and method for public transportation. Instead of modeling passenger attribute recognition as a forced classification problem of "input image to specific attribute category," this system introduces attribute-level observability prediction, attribute-level uncertainty estimation, and risk-calibrated selective output mechanisms. This enables the system to determine whether each attribute has sufficient visual evidence in the current frame and outputs an unknown state when the risk is high.
[0083] I. Definition of Attribute Sets
[0084] In this invention, the input bus passenger image is I, and the attribute set is:
[0085] A={a1,a2,…,a M} ;
[0086] Where M represents the number of attributes. In one implementation, the attribute set includes gender, age group, clothing type, clothing color, sleeve length, hairstyle, mask, glasses, hat, backpack, bottom type, bottom color, shoes, and carried items, etc. For the i-th attribute, the system simultaneously predicts the attribute category y. i Attribute-level observability v i Attribute-level uncertainty u i and reliability decision r i Among them, observable states include at least three categories: visible, partially visible, and invisible.
[0087] II. System Module Architecture
[0088] The present invention provides a selective identification system for public transport passenger attributes, such as... Figure 1 and Figure 2 As shown, it includes an image acquisition and preprocessing module, an attribute recognition and observability output module, an uncertainty estimation unit, a risk calculation and calibration unit, and a selective output unit.
[0089] 1. Image Acquisition and Preprocessing Module
[0090] In one specific embodiment of the present invention, the image acquisition preprocessing module is used to acquire bus cabin surveillance video images. Since bus cabin surveillance images typically contain multiple passengers, the image acquisition preprocessing module determines the passenger target area corresponding to each passenger from the bus cabin surveillance image, and generates a passenger image to be identified based on the passenger target area.
[0091] In one implementation, the image acquisition preprocessing module can extract video frames from the bus monitoring video and use a target detection algorithm to detect passenger targets from the video frames to obtain passenger target boxes; then, based on the passenger target boxes, it can extract the corresponding passenger image regions from the video frames and perform one or more of the following processes on the extracted passenger image regions: size adjustment, pixel normalization, and format conversion, to generate a passenger image to be identified, which serves as the input to the subsequent attribute recognition and observability output module.
[0092] In one exemplary embodiment, the size of the passenger image to be identified can be normalized to (224*224) pixels. It should be noted that the present invention does not specifically limit the method for determining the passenger target region; the passenger target region can be determined by a target detection model, target tracking results, instance segmentation results, a target region provided by an external system, or a manually specified region.
[0093] 2. Attribute Recognition and Observability Output Module
[0094] The attribute recognition and observability output module is used to receive the passenger image to be identified and output the attribute category probability distribution and predicted category, as well as the attribute-level observable state probability vector for each passenger attribute.
[0095] The attribute recognition and observability output module includes a feature extraction unit, a coarse-grained region pooling unit, an attribute classification branch, and an attribute-level observability prediction branch.
[0096] 2.1 Feature Extraction Unit
[0097] The feature extraction unit takes the passenger image I to be identified as input and obtains a shared feature map. :
[0098] (1)
[0099] in, The parameter is Feature extraction network, Let C represent the set of real numbers, and C be the number of feature channels. and These represent the feature map height and width, respectively. The feature extraction unit can use MobileNetV3-Large, ResNet50, ConvNeXt-Tiny, ShuffleNetV2, or other convolutional neural network architectures.
[0100] 2.2 Coarse-grained regional pooling unit
[0101] Coarse-grained regional pooling units share feature maps Coarse-grained region pooling is performed to obtain regional features related to the passenger's body area.
[0102] For shared feature maps Coarse-grained region segmentation is performed along the vertical direction to obtain regional features related to the passenger's body area. Let the feature map height be... In one implementation, it is divided into four regions: a head / face region, an upper body region, a lower body region, and a foot region. For example, the head / face region corresponds to 0%–20% of the feature map height, the upper body region to 20%–50%, the lower body region to 50%–80%, and the foot region to 80%–100%. For the first… There are several regions, with an initial proportion of [missing information]. The completion rate is The corresponding feature map height range is:
[0103] , (2)
[0104] Perform a pooling operation on this region to obtain the region's feature vector:
[0105] (3)
[0106] Wherein, Pool can be average pooling, max pooling, or a combination of both. In a preferred embodiment, global average pooling is used, resulting in:
[0107] (4)
[0108] in, GAP stands for Global Average Pooling.
[0109] This coarse-grained region pooling method has low computational overhead, making it suitable for real-time processing in bus monitoring. Its technical role is to provide body region priors for different attributes: masks, glasses, and hairstyles primarily rely on the head / face region; clothing type, color, and sleeve length primarily rely on the upper body region; bottom type and color primarily rely on the lower body region; shoes primarily rely on the foot region; backpacks and carried items can rely on the upper body region and the area near the hands. Therefore, this region pooling module can assist the attribute classification branch and attribute-level observability branch in determining whether a specific attribute has sufficient visual evidence.
[0110] Let the global feature be:
[0111] (5)
[0112] The global features and features from each region can then be fused to obtain fused features for attribute prediction:
[0113] h (6)
[0114] in, This indicates feature concatenation. In addition to concatenation, weighted summation, attention fusion, or other feature fusion methods can also be used.
[0115] It should be noted that the coarse-grained region pooling described above is only one optional implementation method, and this invention does not depend on any specific region partitioning method. In implementations that do not employ region pooling, the attribute classification branch and the observability prediction branch can also be directly based on the globally shared feature map. Make predictions.
[0116] 2.3 Attribute Classification Branches
[0117] The attribute classification branch is based on the shared feature map. Based on the aforementioned regional features, the attribute category probability distribution and predicted category are output for each passenger attribute. For the i-th passenger attribute, the attribute classification branch outputs the attribute category probability distribution:
[0118] (7)
[0119] in, This represents the number of candidate attribute categories contained in the i-th passenger attribute.
[0120] Classification confidence of the i-th passenger attribute for:
[0121] (8)
[0122] The predicted category for the i-th passenger attribute is:
[0123] (9)
[0124] In one implementation, the attribute classification branch can consist of one or more fully connected layers, and may also include non-linear activation functions, normalization layers, and Dropout layers. Each attribute can have an independent classification header, or multiple attributes can share a portion of the classification layer before being output separately.
[0125] 2.4 Attribute-level observability prediction branch
[0126] The attribute-level observability prediction branch is used based on the shared feature map. Based on the aforementioned regional features, attribute-level observable state probability vectors are output for each passenger attribute.
[0127] Unlike traditional image-level occlusion assessment, this invention predicts the observability state of each attribute in the current image separately. For the i-th attribute, the attribute-level observability prediction branch outputs an observable state prediction score vector:
[0128] (10)
[0129] The Softmax function is then used to convert the observable state prediction score vector into an attribute-level observable state probability vector.
[0130] (11)
[0131] And satisfy:
[0132] (12)
[0133] in, , and These represent the probabilities of the i-th passenger attribute being in a visible, partially visible, and invisible state, respectively. A visible state indicates that the visual area corresponding to the passenger attribute is clearly visible; a partially visible state indicates that the visual area corresponding to the passenger attribute is partially visible or partially occluded; and an invisible state indicates that the visual area corresponding to the passenger attribute is invisible or lacks reliable visual evidence.
[0134] The invisibility probability of the i-th passenger attribute is the probability component of the corresponding invisible state in the observable state probability vector of the attribute level:
[0135] (13)
[0136] in, This represents the observable state of the i-th passenger attribute. This indicates the image of the passenger to be identified.
[0137] The attribute-level observability prediction branch shares the low-level visual features extracted by the backbone network with the attribute classification branch, and can simultaneously utilize region features obtained from coarse-grained region pooling. This branch performs supervised learning through attribute-level observability labels, enabling the system to independently determine whether each attribute possesses reliable visual evidence.
[0138] 3. Uncertainty Estimation Unit
[0139] The uncertainty estimation unit calculates the attribute-level uncertainty based on the attribute category probability distribution output by the attribute classification branch. Attribute-level uncertainty is used to represent the degree of uncertainty of the model's prediction result for the i-th attribute category. The larger the value, the more uncertain the model is about predicting the attribute category. In one implementation, attribute-level uncertainty... Calculate according to the following formula:
[0140] (14)
[0141] in, This represents the classification confidence score of the i-th passenger attribute. In other embodiments, prediction entropy can be used to measure attribute-level uncertainty, or normalized prediction entropy, Monte Carlo Dropout, model ensemble, or Bayesian neural networks can be used to obtain attribute-level uncertainty. This invention does not limit the specific calculation method of attribute-level uncertainty, as long as it can reflect the degree of uncertainty of the model's prediction result for passenger attribute categories.
[0142] 4. Risk Calculation and Calibration Unit
[0143] The risk calculation and calibration unit is used to calculate the invisibility probability of the i-th passenger attribute. and attribute level uncertainty Merge into attribute-level risk scores :
[0144] (15)
[0145] in, and These represent the risk weights for the probability of an attribute being invisible and the uncertainty, respectively. In a preferred embodiment, a configuration is provided. This ensures that the risk score is primarily influenced by the invisible probability of attributes, while also taking into account classification uncertainty.
[0146] Risk Score The technical meaning is: if the current system outputs the specific category of the i-th attribute The degree of risk that the output lacks support from current visual evidence of the image. The higher the value, the more likely the local area corresponding to the attribute is to be invisible; The higher the threshold, the less evidence there is for decision-making within the attribute classification branch. The risk calculation and calibration unit also determines the attribute-level risk calibration threshold for each passenger attribute based on the validation set. .
[0147] Calibration threshold The technical meaning is that the system determines reliable output boundaries for different attributes. For example, attributes such as shoes, masks, and glasses are affected by occlusion to different degrees, so their risk thresholds can be calibrated separately to achieve fine control over the risk of erroneous output for different attributes.
[0148] In one implementation, the attribute-level calibration threshold is determined by the following optimization objective:
[0149] (16)
[0150] in, Represents the set of candidate thresholds Search for the threshold that maximizes the expression; For the target risk level (e.g.) =0.05、 =0.10 or =0.20).
[0151] Risk of the i-th attribute being invisible on the validation set:
[0152] (17)
[0153] The coverage of the i-th attribute as visible / partially visible on the validation set:
[0154] = (18)
[0155] in, Indicates the number of samples in the validation set. This represents the true observable state label of the i-th passenger attribute in the n-th validation sample. This indicates the corresponding attribute-level risk score. Represents the set of candidate thresholds. This represents the upper limit of the false output rate for a specific category of the invisible attribute. 1 (⋅) is the indicator function, which takes the value of 1 when the condition in parentheses is true, and takes the value of 0 otherwise.
[0156] That is, the risk of erroneous output in a specific category of invisible attributes does not exceed the target risk level. Under the premise of maximizing the output coverage of visible and partially visible attributes, the attribute-level risk calibration threshold can be determined as much as possible. In other implementations, the attribute-level risk calibration threshold can also be determined using the equal error rate criterion, Bayesian decision criterion, or operation point selection method based on the risk-coverage curve.
[0157] 5. Selective Output Unit
[0158] The selective output unit is used to determine the attribute level risk score based on the i-th passenger attribute. Corresponding attribute-level risk calibration threshold The comparison results will output the specific attribute category or unknown status.
[0159] If the risk score of the i-th passenger attribute satisfies If the attribute classification branch outputs a specific attribute category, then the attribute is considered a reliable output, and the predicted category is output; if... If the condition is not met, the attribute is determined to be an unreliable output, and the output state is unknown. The final selective attribute output is:
[0160] (19)
[0161] Through the aforementioned selective output mechanism, the system can output specific attribute categories for passenger attributes with sufficient visual evidence and low classification risk, and output unknown status for passenger attributes with high risk or lack of reliable visual evidence. This avoids forcibly outputting specific attribute categories when reliable visual evidence is lacking and reduces the rate at which truly invisible attributes are incorrectly output as specific categories.
[0162] III. Dataset Construction and Annotation
[0163] In this embodiment, the training dataset is derived from bus cabin surveillance video frames. After acquiring video images from the bus cabin monitoring equipment, passenger target regions are obtained using object detectors (such as YOLO series, Faster R-CNN, or other pedestrian detection models) or manual annotation. Passenger targets are cropped based on the detection boxes, and the cropped passenger images are normalized to a uniform size (e.g., 224×224 pixels) to obtain passenger cropped image I. For each passenger sample, two types of information are annotated:
[0164] The first category is attribute category labels, used to describe the semantic attributes of passengers. In one implementation, the attribute set includes gender, age group, clothing type, clothing color, sleeve length, hairstyle, mask, glasses, hat, backpack, bottom type, bottom color, shoes, and carried items. The category value of each attribute can be defined according to the actual application scenario. For example, the mask attribute may include "with mask" and "without mask," and the clothing color attribute may include categories such as "black," "white," and "blue."
[0165] The second category is attribute-level observability labels, used to describe whether each attribute has reliable visual evidence in the current image. Attribute-level observability labels include three categories: visible (vis), partially visible (par), and invisible (inv). Visible indicates that the region corresponding to the attribute is clearly visible and has sufficient visual evidence; partially visible indicates that the region corresponding to the attribute is partially visible or slightly occluded, but still provides weak supervision information; invisible indicates that the region corresponding to the attribute is completely invisible or lacks reliable visual evidence.
[0166] It should be noted that the observability of different attributes can vary within the same passenger image. For example, in an image of a bus, the upper body of a passenger may be clearly visible, so the color of the shirt can be labeled as visible; however, the passenger's feet may be obscured by the seat, so the shoe attribute can be labeled as invisible. Therefore, this invention employs attribute-level observability labeling, rather than solely relying on image-level occlusion labeling.
[0167] The labeled dataset is divided into training, validation, and test sets according to a preset ratio. For example, 80% of the images can be used as the training set, 10% as the validation set, and 10% as the test set; alternatively, a ratio of 70% training, 15% validation, 15% test, or other ratios can be used. The specific ratio can be adjusted based on the total amount of data and the model complexity. The training set is used for learning model parameters, the validation set is used for determining attribute-level risk calibration thresholds, and the test set is used for testing model performance.
[0168] It should be noted that the present invention does not limit the source, size, or specific division ratio of the dataset; as long as it contains the above-mentioned tag information, the technical solution of the present invention can be achieved.
[0169] IV. Network Model Training
[0170] In this embodiment, the attribute classification branch and the attribute-level observability prediction branch are jointly trained in the following manner.
[0171] First, training images of bus passengers are acquired, each labeled with an attribute category and an attribute-level observability label. The attribute-level observability labels include visible, partially visible, and invisible states.
[0172] For the The training image of the th ... Each passenger attribute is labeled with its actual attribute level and observability. Set the attribute classification loss weights:
[0173] (20)
[0174] Where 0 < α < 1. In a preferred embodiment, α = 0.5. This weight setting is used to ensure that visible attributes fully participate in attribute category supervision, some visible attributes participate in attribute category supervision with lower weights, and invisible attributes do not participate in specific category supervision, thereby avoiding the model learning unfounded category associations on attributes lacking visual evidence.
[0175] Calculate the observability-weighted attribute classification loss based on the attribute classification loss weights:
[0176] ; (twenty one)
[0177] in, Indicates attribute category label, Represents the probability distribution of attribute categories. Represents the cross-entropy loss function. To prevent small constants with a denominator of zero;
[0178] Calculate the observability prediction loss based on the attribute-level observability labels:
[0179] ; (twenty two)
[0180] in, Let A represent the probability vector of observable states at the attribute level, where N represents the number of training images and A represents the number of passenger attributes.
[0181] Finally, based on the observability-weighted attribute classification loss and the observability prediction loss, the attribute classification branch and the attribute-level observability prediction branch are jointly trained:
[0182] L + ; (twenty three)
[0183] in, The weights represent the observability prediction loss. In one implementation, The value can be between 0.2 and 0.4, with 0.4 being the preferred value.
[0184] In one implementation, model training uses the Adam or AdamW optimizer. The initial learning rate can be set to...
[0185] 1×10 −4 ~1×10 −3 The weight decay can be set to 1×10. −4 The batch size can be set to 32 or 64 MB depending on the device's video memory, and the total number of training epochs can be set to 50 to 100. The learning rate can employ a cosine annealing decay strategy, a step decay strategy, or other commonly used learning rate adjustment strategies. The above training hyperparameters can be adjusted according to the data scale, backbone network structure, and deployment requirements, and do not constitute a limitation on the scope of protection of this invention.
[0186] Through the joint training described above, the attribute-level observability prediction branch can learn to determine the visible, partially visible, and invisible states of each passenger attribute, while the attribute classification branch can learn specific category discrimination ability on attributes with reliable visual evidence and reduce the learning of category associations without reliable visual basis on invisible attributes.
[0187] V. Experimental Verification
[0188] To verify the effectiveness of this invention, a comparative experiment was conducted on a bus passenger compartment surveillance image dataset. The experiment compared the performance of standard attribute recognition methods, PAR+unknown category methods, visibility mask PAR methods, OA-PAR fixed gating methods, confidence threshold methods, visibility-only gating methods, and the method of this invention in terms of the false output rate of specific categories of invisible attributes.
[0189] Experimental results are as follows Figure 4As shown. Standard attribute identification methods, due to their forced output of specific categories for all attributes, have a high false output rate for specific categories on invisible attributes. The PAR+unknown category method and the OA-PAR fixed gating method can reduce the false output rate to some extent. The confidence threshold method and the visibility-only gating method can further reduce the false output rate of specific categories for invisible attributes. The method of this invention, by fusing attribute-level invisibility probability and attribute-level uncertainty, and using attribute-level risk calibration thresholds for selective output, can reduce the false output rate of specific categories for invisible attributes to approximately 8.6% and 6.4% at target risk levels ρ=0.10 and ρ=0.05, respectively.
[0190] To evaluate the effectiveness of this invention in suppressing erroneous output of invisible attributes, the risk of output for specific categories of invisible attributes is defined as follows:
[0191] ; (twenty four)
[0192] This metric measures the proportion of truly invisible attributes that are still output as specific categories by the system. The lower the value, the lower the risk of the system mis-outputting invisible attributes.
[0193] The experimental results above show that by introducing attribute-level observability prediction, attribute-level uncertainty estimation, risk calculation and calibration, and selective output mechanism, this invention can reduce the rate at which invisible attributes in bus cabin monitoring images are incorrectly output as specific categories, thereby improving the reliability of bus passenger attribute recognition results.
[0194] VI. Computer-readable storage media
[0195] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0196] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications, equivalent substitutions, and improvements can be made to the present invention without departing from its spirit and principles, and all such modifications, equivalent substitutions, and improvements should be included within the scope of protection of the present invention.
Claims
1. A selective identification system for bus passenger attributes, characterized in that: It includes an image acquisition and preprocessing module, an attribute recognition and observability output module, an uncertainty estimation unit, a risk calculation and calibration unit, and a selective output unit; The image acquisition and preprocessing module obtains a single passenger image; Used to acquire bus surveillance video images, determine passenger target areas from the bus surveillance video images, and generate passenger images to be identified based on the passenger target areas; The attribute recognition and observability output module outputs the attribute prediction category and the attribute-level observable state probability vector, respectively. This is used to receive the image of the passenger to be identified, and to output the attribute category probability distribution and predicted category, as well as the attribute-level observable state probability vector, for each passenger attribute; wherein, for the i-th passenger attribute, the attribute category probability distribution is expressed as... The attribute-level observable state probability vector includes the visibility probability. Partially visible probability and invisibility probability "Visible" indicates that the visual area corresponding to the attribute is clearly visible; "partially visible" indicates that the area corresponding to the attribute is partially visible or partially occluded; "invisible" indicates that the area corresponding to the attribute is invisible or lacks reliable visual evidence. The uncertainty estimation unit calculates attribute-level uncertainty; Based on the attribute category probability distribution of the i-th passenger attribute Computational property level uncertainty ; The probability of feature attributes being invisible and attribute level uncertainty Merge into attribute-level risk scores : ; in, and These represent the risk weights for the probability of an attribute being invisible and the uncertainty, respectively. The risk calculation and calibration unit determines the attribute-level risk score and calibration threshold of the characteristic attribute; based on the attribute-level uncertainty... and invisibility probability Determine the attribute level risk score s for the i-th passenger attribute. i And based on the validation set, learn the attribute-level risk calibration threshold for each feature attribute. ; The selective output unit is used to output a specific attribute category or an unknown state; If the risk score of the i-th feature attribute satisfies If the attribute classification branch outputs a specific attribute category, then the attribute is considered a reliable output, and the predicted category is output. ; like If the attribute is determined to lack reliable visual evidence, then the attribute is determined to be an unreliable output, and the output state is unknown.
2. The selective identification system for bus passenger attributes according to claim 1, characterized in that: The risk calculation and calibration unit is used to determine the corresponding attribute-level risk calibration threshold for the i-th passenger attribute based on the attribute-level risk score in the validation set and the actual attribute-level observable state label. : The risk calculation and calibration unit calculates the false output rate of the specific category of the invisible attribute for the i-th passenger attribute based on the number of samples in the verification set whose attribute-level risk scores are not greater than the candidate threshold τ. ; The risk calculation and calibration unit calculates the effective attribute coverage of the i-th passenger attribute based on the number of samples in the verification set whose attribute-level risk scores are not greater than the candidate threshold τ, both in the truly visible samples and the truly partially visible samples. = ; The risk calculation and calibration unit ensures that the false output rate of the specific category of the invisible attribute is not greater than the preset upper limit of the false output rate of the specific category of the invisible attribute. Under the condition of, from the candidate threshold set of the i-th passenger attribute Select the candidate threshold that maximizes the coverage of the effective attributes, and use it as the attribute-level risk calibration threshold for the i-th passenger attribute: ; in, Indicates the number of samples in the validation set. This represents the true observable state label of the i-th passenger attribute in the n-th validation sample. This indicates the corresponding attribute-level risk score. Represents the set of candidate thresholds. This represents the upper limit of the error output rate for a specific category of the preset invisible attribute, and 1(⋅) is the indicator function.
3. The selective identification system for bus passenger attributes according to claim 1, characterized in that: The uncertainty estimation unit is used to estimate the probability distribution of the attribute category of the i-th passenger attribute. Computational property level uncertainty : Wherein, the attribute category probability distribution of the i-th passenger attribute Represented as: ; Classification confidence of the i-th passenger attribute for: ; The attribute level uncertainty Calculate according to the following formula: ; in, This represents the number of candidate attribute categories for the i-th passenger attribute. This represents the probability that the i-th passenger attribute belongs to the k-th candidate attribute category. The larger the value, the more uncertain the model's prediction of the category of the i-th passenger attribute.
4. The selective identification system for bus passenger attributes according to claim 1, characterized in that: The image acquisition and preprocessing module includes an image acquisition unit and a passenger target region generation unit; The image acquisition unit is used to acquire monitoring images of the bus compartment; The passenger target area generation unit is used to determine at least one passenger target area from the bus carriage monitoring image, and generate a corresponding passenger image to be identified based on the passenger target area.
5. The selective identification system for bus passenger attributes according to claim 1, characterized in that: The attribute recognition and observability output module includes a feature extraction unit, a coarse-grained region pooling unit, an attribute classification branch, and an attribute-level observability prediction branch. The feature extraction unit is used to receive the passenger image to be identified and extract the shared feature map F from the passenger image to be identified; The coarse-grained region pooling unit is used to perform coarse-grained region pooling on the shared feature map F to obtain the region features corresponding to the passenger's body region; The attribute classification branch is used to output the attribute category probability distribution and predicted category for each passenger attribute based on the shared feature map F and the region features; The attribute-level observability prediction branch, based on shared features and regional features, is used to output attribute-level observable state probability vectors for each passenger attribute based on the shared feature map F and the regional features; wherein, for the i-th passenger attribute, the attribute-level observable state probability vector includes the visibility probability. Partially visible probability and invisibility probability .
6. The selective identification system for bus passenger attributes according to claim 5, characterized in that: The attribute-level observability prediction branch is used to predict the attribute-level observability state of each passenger attribute in the current image of the passenger to be identified. For the i-th attribute, the attribute-level observability prediction branch outputs an observable state prediction score vector: ; The Softmax function is then used to convert the observable state prediction score vector into an attribute-level observable state probability vector. ; in, , and These represent the probabilities of the i-th passenger attribute being in a visible, partially visible, and invisible state, respectively. The visible state indicates that the visual area corresponding to the passenger attribute is clearly visible, the partially visible state indicates that the visual area corresponding to the passenger attribute is partially visible or partially occluded, and the invisible state indicates that the visual area corresponding to the passenger attribute is invisible or lacks reliable visual evidence. The invisibility probability of the i-th passenger attribute is the probability component of the corresponding invisible state in the observable state probability vector of the attribute level: ; in, This represents the observable state of the i-th passenger attribute. This indicates the image of the passenger to be identified.
7. A method for selectively identifying passenger attributes on public transportation, characterized in that, Includes the following steps: 1) Acquire bus cabin surveillance images, determine passenger target areas from the bus cabin surveillance images, and generate passenger images to be identified based on the passenger target areas; 2) Extract features from the image of the passenger to be identified to obtain a shared feature map F; perform coarse-grained region pooling on the shared feature map F to obtain region features related to the passenger's body area; The attribute classification branch outputs the predicted category and confidence score for each passenger attribute; the observability prediction branch outputs the visibility probability, partial visibility probability, and invisibility probability for each passenger attribute, where the invisibility probability is denoted as... ; 3) Calculate the attribute level uncertainty based on the attribute category probability distribution of the passenger attributes. The probability of the attribute being invisible. and attribute level uncertainty Merge into attribute-level risk scores : ; in, and These represent the risk weights for the probability of an attribute being invisible and the uncertainty, respectively. 4) Determine the attribute-level risk calibration threshold for each passenger attribute based on the validation set. ; 5) If the risk score of the i-th passenger attribute satisfies If the attribute classification branch outputs a specific attribute category, then the attribute is considered a reliable output, and the predicted category is output. ; like If the attribute is determined to lack reliable visual evidence, then the attribute is determined to be an unreliable output, and the output state is unknown.
8. The method for selectively identifying passenger attributes on public transport according to claim 7, characterized in that, The attribute recognition and observability output module includes a feature extraction unit, a coarse-grained region pooling unit, an attribute classification branch, and an attribute-level observability prediction branch. Step 2) includes the following steps: 2.1) The feature extraction unit takes the passenger image to be identified as input and obtains a shared feature map F; 2.2) The coarse-grained region pooling unit performs coarse-grained region pooling on the shared feature map F to obtain region features related to the passenger's body region; 2.3) The attribute classification branch, based on the shared feature F and the regional feature, outputs the category probability distribution and predicted category for each passenger attribute; 2.4) The attribute-level observability prediction branch outputs the visibility probability, partial visibility probability, and invisibility probability of each passenger attribute based on the shared feature F and the regional feature.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the selective identification method for bus passenger attributes as described in claim 7 or 8.