Vehicle passenger number estimation method based on distributed acoustic sensing time sequence signals

By combining a feature extraction backbone network of one-dimensional convolution and Transformer with self-supervised pre-training of a mask autoencoder, the problems of low accuracy and poor robustness in vehicle occupant number recognition in existing technologies are solved, achieving high-precision and low-cost vehicle occupant number recognition, which is suitable for multiple application scenarios.

CN121786622APending Publication Date: 2026-04-03ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing vehicle identification methods based on distributed acoustic sensing have low accuracy in identifying the number of occupants inside the vehicle, poor robustness to conditions with a small number of occupants, strong dependence on labeled data, and do not fully utilize the potential information of unlabeled data. They are difficult to meet the high accuracy and high robustness requirements of occupant identification in carpool lane supervision and differentiated scenarios.

Method used

A feature extraction backbone network combining one-dimensional convolution and Transformer is adopted, and self-supervised pre-training is performed by combining masked autoencoder. General temporal features are learned using unlabeled DAS data, and the model is constructed by supervising fine-tuning training with label smoothing to realize the automatic identification of vehicle occupancy rate.

Benefits of technology

It improves the accuracy and robustness of occupant number recognition, reduces reliance on labeled data, has inherent privacy protection features, facilitates large-scale deployment and maintenance, adapts to different scenarios and operating conditions, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786622A_ABST
    Figure CN121786622A_ABST
Patent Text Reader

Abstract

A vehicle passenger number estimation method based on a distributed acoustic sensing time sequence signal comprises the following steps: S1, obtaining an original one-dimensional vibration signal generated when a vehicle passes by, and preprocessing the original one-dimensional vibration signal to form a one-dimensional time sequence sample library with an in-vehicle occupancy rate identifier; s2, constructing a one-dimensional convolution and Transform combined feature extraction backbone network, performing layered down-sampling and feature extraction on the long-sequence one-dimensional DAS signal, and changing an original long-sequence vibration signal into a plurality of feature token sequences with shorter lengths; s3, taking the backbone network as an encoder and additionally arranging a decoding module to form a complete mask auto-encoder (MAE) under the condition that no tag is needed, performing a sequence reconstruction task by utilizing an MAE decoder after random masking to mine general time sequence characteristics in a large amount of non-tag DAS data to optimize parameters of the MAE encoder, and obtaining a complete mask auto-encoder (MAE); self-supervised pre-training is carried out on a preprocessed DAS time sequence sample, and general time sequence feature representation related to vehicle passing is learned and obtained by means of a random mask and a reconstruction task; s4, based on the DAS data set with the in-vehicle passenger number label, removing an MAE decoder, only retaining a pre-trained MAE encoder, adding a classification head at the output tail end of the MAE encoder, and executing supervision fine tuning training to obtain a multi-classification model specially used for identifying the in-vehicle occupancy rate; and S5, taking a new DAS time sequence obtained online as input to the in-vehicle occupancy rate classification model trained in the step S4, and giving an output result of the type of the number of passengers in the vehicle for each time window so as to achieve the purpose of automatically identifying the in-vehicle occupancy rate of the passing vehicle. According to the method, a self-supervised pre-training mechanism is introduced, so that the problems that a traditional method depends on a large amount of labeled data and is sensitive to specific acquisition conditions are solved, real-time and non-contact intelligent detection of the occupancy rate in the vehicle is realized, and meanwhile, the method has good privacy protection capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of intelligent transportation and distributed optical fiber sensing, specifically relating to a means and system for identifying the number of occupants in a vehicle using distributed acoustic sensing (DAS), with a particular focus on a method for determining the number of occupants in a single vehicle by using a one-dimensional mask autoencoder for self-supervised pre-training. Background Technology

[0002] With the development of intelligent transportation systems (ITS), accurate perception of road condition monitoring, traffic flow statistics, and lane occupancy has become a crucial foundation for traffic safety management and optimal allocation of road resources. Currently, commonly used sensing methods include video surveillance, inductive loop detectors, and millimeter-wave radar. These solutions typically require the deployment of dedicated hardware on or along roadsides, resulting in significant construction disruptions, high maintenance costs, and poor environmental adaptability. A key focus of ITS is encouraging multi-occupant vehicles, reducing the time spent in these vehicles relative to passenger vehicles, thereby decreasing carbon dioxide emissions and fuel consumption. Therefore, vehicle occupancy detection is paramount. Existing methods for identifying the number of occupants in a vehicle primarily rely on onboard cameras or roadside high-definition cameras, analyzing in-vehicle images through image processing or deep learning methods. These methods have the following problems: First, they require the collection of occupant facial or in-vehicle images, which poses a significant privacy risk; second, the recognition effect is highly sensitive to lighting conditions and weather conditions, and the recognition rate will decrease in environments such as nighttime, backlight, rain, and fog; third, the installation and maintenance costs of roadside cameras and supporting networks are high, making it difficult to deploy them on a large scale over long distances.

[0003] Distributed acoustic sensing (DAS) utilizes ordinary optical fiber as the sensing medium, and through back-end optoelectronic equipment and signal processing units, achieves high-sensitivity acquisition of vibration signals along the road. Laying optical fiber near roads or bridges allows for continuous acquisition of vibration signals generated by passing vehicles without damaging the road surface. This data is used for vehicle detection, classification, and occupancy identification, offering advantages such as flexible deployment, low maintenance costs, and wide coverage. Existing DAS-based traffic applications primarily focus on vehicle detection, speed estimation, lane recognition, and axle count identification, with relatively little research on occupant identification. Furthermore, existing DAS-based vehicle occupancy identification methods often employ direct training of one-dimensional convolutional neural networks (1D-CNN) or other deep learning models on labeled DAS time-series signals. This type of method suffers from the following problems: 1. Limited and unevenly distributed labeled data: In real-world engineering, the amount of data that can accurately label the number of passengers in vehicles and lane information is limited. Moreover, the sample distribution of the number of passengers in different vehicles is uneven, with fewer samples having a high occupancy rate. This makes it easy for direct supervised training to lead to overfitting and class bias.

[0004] 2. There is a significant discrepancy between training data and actual conditions: Training data is usually collected using a specific method, but actual application scenarios may involve different road sections and different paving methods. In this case, even if traditional 1D-CNN achieves a high accuracy rate on the training data, its recognition accuracy in a single working condition may not meet expectations.

[0005] 3. Low utilization of large amounts of unlabeled DAS data: A large amount of unlabeled DAS time-series data can be obtained in the field, containing rich information on road vibration patterns and environmental noise. However, many existing methods often rely on a small amount of labeled data to perform supervised learning, which is difficult to obtain labels and does not make good use of the temporal characteristics hidden in unlabeled data, thus limiting the generalization ability of the model.

[0006] In summary, existing methods mostly train 1D-CNNs directly on DAS data with limited annotations, failing to fully utilize the potential information from large amounts of unlabeled data. They also lack cost-sensitive design for challenging occupant count categories, resulting in insufficient recognition accuracy in independent scenarios. This makes it difficult to meet the high accuracy and robustness requirements of vehicle occupant count recognition in carpooling lane monitoring and differentiated scenarios. Therefore, research on vehicle occupant count recognition methods remains essential. Summary of the Invention

[0007] To overcome the problems of existing vehicle identification methods based on distributed acoustic sensing, such as low accuracy in identifying the number of occupants inside the vehicle, poor robustness to conditions with a small number of occupants, and strong dependence on labeled data, this invention provides a method for identifying the number of occupants inside the vehicle based on distributed acoustic sensing time-series signals.

[0008] The present invention provides a method for identifying the number of occupants inside a vehicle based on distributed acoustic sensing time-series signals, the method comprising the following steps: S1: A distributed fiber optic acoustic sensing system laid under the road is used to obtain the original one-dimensional vibration signal generated when a vehicle passes by. Channel selection, window truncation and marking are performed on the original fiber optic data to form a one-dimensional time series sample library with vehicle occupancy indicators.

[0009] S2: Construct a feature extraction backbone network combining one-dimensional convolution and Transformer to perform hierarchical downsampling and feature extraction on long-sequence one-dimensional DAS signals, transforming the original long-sequence vibration signal into several shorter feature token sequences, providing a consistent backbone feature space for subsequent MAE pre-training and supervised fine-tuning.

[0010] S3: Without labels, the backbone network constructed in step S2 is used as an encoder, and a decoding module is added to form a complete masked autoencoder (MAE). The MAE decoder is used to perform a sequence reconstruction task after random masking to mine general temporal features in a large amount of unlabeled DAS data and optimize the MAE encoder parameters. Self-supervised pre-training is performed on the DAS time series samples preprocessed in S1. The random mask and reconstruction task are used to learn and obtain general temporal feature representations of vehicle passage.

[0011] S4: Based on the DAS dataset with labels for the number of occupants in the vehicle, the MAE decoder is removed, and only the pre-trained one-dimensional convolutional + Transformer backbone network (i.e., the MAE encoder) is retained. A classification head is added to the end of its output, and supervised fine-tuning training is performed using label smoothing to obtain a multi-classification model specifically for recognizing vehicle occupancy.

[0012] S5: Use the newly acquired DAS time series as input to the vehicle occupancy classification model trained in step S4, and output the number of occupants in the vehicle for each time window to achieve the purpose of automatic identification of the vehicle occupancy rate of passing vehicles.

[0013] Step S1 specifically includes: using a distributed optical fiber sensing system laid under the road to collect multi-channel one-dimensional vibration signals of the target road segment during vehicle passage, and saving the collected raw data in the form of a time-series array. Based on the known vehicle lane position and driving direction, several continuous spatial channels covering the vehicle trajectory are selected as the range of channels of interest, and the raw DAS signal is windowed on the time axis according to the start and end times of the vehicle passage event to obtain a fixed-length one-dimensional time series segment.

[0014] Furthermore, by combining the vehicle number, driving speed, driving direction, and number of occupants recorded in the experimental conditions, each time series segment was paired with a vehicle, and the samples were classified into the corresponding occupancy category according to the number of occupants. This constructed a training sample set with the DAS one-dimensional time series as input and the occupant number category as the output label. Zero-padding or padding was performed on sequence segments with insufficient time length to ensure that all samples maintained consistent length in the time dimension, thus adapting to the input requirements of subsequent neural network training.

[0015] Step S2 specifically includes: based on the labeled DAS one-dimensional time series obtained in step S1, constructing a feature extraction backbone network consisting of a multi-layer one-dimensional convolutional network and a Transformer encoder connected in series; firstly, using several levels of one-dimensional convolutional layers in conjunction with pooling or strided convolution operations, performing hierarchical downsampling on the original long sequence. Each one-dimensional convolutional layer extracts local vibration patterns in the time direction, and gradually shortens the sequence length and increases the number of channels, thereby suppressing high-frequency noise while preserving the key signal of the vehicle passing through. After multiple downsampling steps, the original length is reduced to... The DAS sequence is compressed to a length of The intermediate feature sequence is used, where the feature vectors at each time point are treated as a set of low-dimensional feature tokens. Then, this intermediate feature sequence is input into the Transformer encoder, which has a length of [missing information]. Multi-head self-attention operations and feedforward network transformation operations are performed on the token sequence to model the long-distance dependencies between different time positions and the overall energy distribution characteristics. By utilizing the self-attention mechanism, the model can simultaneously capture the comprehensive features of key time periods such as vehicle entry, passage, and departure from the sensing area, and obtain feature token representations containing overall vehicle dynamic information and occupancy-related patterns, thus providing a consistent backbone feature space for subsequent self-supervised pre-training and supervised classification.

[0016] Step S3 specifically includes: Based on the one-dimensional convolutional and Transformer backbone network described in step S2, a time-scale decoding module is added, which, together with the backbone network, constitutes a one-dimensional masked autoencoder (MAE) structure for self-supervised pre-training of unlabeled DAS time series. First, each DAS time series is cut into continuous time patches of a fixed length. Then, according to a preset masking ratio, most patches are randomly selected for masking, with only a small number of patches retained as visible tokens. The masked patches are replaced with special mask markers. Next, the masked time series is input into the one-dimensional convolutional + Transformer encoder of step S2 to obtain a compressed latent space feature representation. Then, the decoding end performs upsampling operations layer by layer to restore the time dimension, thereby reconstructing the complete DAS time series segment. Finally, the mean squared error between the reconstructed sequence and the original unmasked sequence is used as the reconstruction loss function to optimize and adjust the parameters of the masked autoencoder.

[0017] Step S4 specifically includes: After completing the self-supervised pre-training described in step S3, the masked autoencoder is removed, and the encoding backbone network composed of one-dimensional convolution and Transformer is retained as the feature extractor; at the end of the backbone network, a classification head for identifying vehicle occupancy rate is added. This classification head includes: a one-dimensional convolutional or fully connected compression layer for further integrating time-dimensional features and compressing them to a smaller scale, followed by several fully connected layers and non-linear activation functions, and finally, the predicted probability of each occupant number category is output using Softmax; in the supervised fine-tuning stage, the DAS time series samples with occupant number labels constructed in step S1 are used as input, and the model is trained using the cross-entropy loss function, while label smoothing is added. The smoothing strategy applies appropriate softening to the original one-hot labels to suppress the model's overconfidence in a single class, reduce the impact of individual sample label noise on the training process, and thus improve the macro-average F1 score, enhancing the robustness of recognition in a few operating conditions such as multiple occupants. During training, by monitoring comprehensive indicators such as macro-average F1 and classification accuracy on the validation set, and using an early stopping strategy to select appropriate training rounds, a vehicle occupancy classification model applicable to various application scenarios and operating conditions with various occupant numbers is obtained.

[0018] Step S5 specifically includes: In the model deployment phase, the new DAS time series acquired online by the distributed fiber optic acoustic wave sensing system is processed according to the same preprocessing procedure as in the training phase, including channel selection, time window truncation, and standardization, and then input into the vehicle occupancy classification model trained in step S4. The model first extracts temporal features through one-dimensional convolution and a Transformer backbone network, and then outputs the predicted probability of each occupant number category by the classification head, and determines the category with the highest probability as the vehicle occupancy rate corresponding to that time window. Based on the vehicle's passage time and trajectory information, the prediction results of multiple windows can be fused to obtain the final occupant number estimate for a single vehicle passage, realizing real-time, non-contact, and image-free intelligent detection of the vehicle occupancy rate.

[0019] The advantages of this invention are: combining a temporal backbone network composed of one-dimensional convolutional neural networks (1D-CNN) and Transformers, it achieves collaborative modeling of large-scale structure and fine-grained features of vehicle vibration / acoustic scattering signals while taking into account both local details and long-range dependencies; using a masked autoencoder (MAE) for self-supervised pre-training, it learns general temporal representations under unlabeled conditions, significantly improving its generalization ability and robustness when facing a small number of samples, class imbalance, and different scenarios; utilizing a "random mask-reconstruction" mechanism to enhance its resistance to noise, missing segments, and changes in operating conditions; employing label smoothing in the supervised fine-tuning stage to suppress overconfident predictions and reduce label noise sensitivity, thereby improving the robustness of macro-average F1 and recognition of operating conditions with a small number of occupants; and, since this scheme relies solely on data from buried distributed optical fiber sensing without collecting image and voice information, it possesses inherent privacy protection characteristics, is conducive to large-scale deployment and easy to maintain, and has engineering feasibility and cost advantages. Attached Figure Description

[0020] Figure 1 This is a flowchart of the in-vehicle occupancy identification method based on distributed optical fiber sensing of the present invention.

[0021] Figure 2 This is a schematic diagram of the structural principle of the convolutional-transformer backbone network of the present invention.

[0022] Figures 3(a) and 3(b) are schematic diagrams of the self-supervised pre-training principle of the masked autoencoder (MAE) of the present invention. Figure 3(a) is a schematic diagram of the one-dimensional time series segmentation, random masking and encoding-decoding process. Figure 3(b) is a schematic diagram of the alignment of the reconstructed output with the original input and the process of calculating the reconstruction loss on the masked segment. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0024] The main process of the vehicle occupant number identification method based on distributed acoustic sensing time-series signals of the present invention is as follows: Figure 1 As shown, it specifically includes: S1: A distributed fiber optic acoustic sensing system laid under the road is used to obtain the original one-dimensional vibration signal generated when a vehicle passes by. Channel selection, window truncation and marking are performed on the original fiber optic data to form a one-dimensional time series sample library with vehicle occupancy indicators.

[0025] S2: Construct a feature extraction backbone network combining one-dimensional convolution and Transformer to perform hierarchical downsampling and feature extraction on long-sequence one-dimensional DAS signals, transforming the original long-sequence vibration signal into several shorter feature token sequences.

[0026] S3: Without labels, the DAS time series samples preprocessed by S1 are self-supervised pre-trained using a masked autoencoder (MAE). The random mask and reconstruction task are used to learn and obtain general temporal feature representations of vehicle passage.

[0027] S4: Based on the DAS dataset with labels of the number of occupants in the vehicle, the MAE decoder is removed, and only the pre-trained one-dimensional convolutional + Transformer backbone network is retained. A classification head is added to the end of its output, and supervised fine-tuning training is performed using label smoothing to obtain a multi-classification model specifically for recognizing vehicle occupancy.

[0028] S5: Use the newly acquired DAS time series as input to the vehicle occupancy classification model trained in step S4, and output the number of occupants in the vehicle for each time window to achieve the purpose of automatic identification of the vehicle occupancy rate of passing vehicles.

[0029] Step S1 specifically includes: for phase difference observations acquired by a distributed optical fiber sensing system (DAS) with optical fibers deployed along the roadside, its discrete matrix on the "spatial channel (bin) × time sampling (shot)" plane is denoted as... Among them, row index Corresponding sampling positions along the fiber length, column index Corresponding to discrete sampling times. Construct a label matrix of the same type. The initial value of this matrix is ​​zero, indicating that it is unlabeled. Because the collected data is extensive, varied, and largely useless, we need to process the data first. The first step is to select the effective channel intervals. It also allows setting selectable noise ranges. Used for rejection (the interval is marked as empty when there is no noise). The set of channels that are validly included in the annotation is defined as: (1) Target occupancy category Its temporal envelope is approximated by two linear boundaries on the "channel-time" plane: (2) These correspond to the start and end time boundaries, respectively. To ensure the index is valid and matches discrete sampling, for each... Calculate the window endpoints after integer and clipping: Take no more than The largest integer and cropped to ; Take not less than Find the smallest integer and crop it to If when This means that the channel is skipped when the window is empty.

[0030] For each retained channel window, assign a category to the sampling points covered by the window in the two-dimensional label plane. The annotation; simultaneously record the window width: (3) To be used in subsequent one-dimensional classification networks, the two-dimensional windows are sequentially stitched together according to the channel size, with the corresponding one-dimensional slices forming the original sequence of the sample: (4) in This indicates a cascaded operation in channel order. This is to meet the fixed input length requirement of the network. It is necessary to Perform length scaling (zero padding if insufficient, symmetrical trimming if excessive) to obtain Finally, the required sample pairs for training / validation are generated. .

[0031] Step S2 specifically includes: For long-term one-dimensional DAS sequences, traditional CNNs rely on a fixed receptive field size to achieve information convergence using convolutional kernels and pooling operations. This type of method, on the one hand, requires "indirectly" expanding the receptive field by stacking multiple layers and increasing the stride, which easily leads to insufficient global dependency characterization and weakened temporal sequence information; on the other hand, when the sequence length is large, it is difficult to establish distant correlations and pattern alignment relationships across different segments based solely on the local inductive bias of convolution itself. Further deepening the network leads to a significant increase in parameters and computational overhead, and a decrease in training stability. Therefore, this step adopts a convolution-transformer structure, adding the ability to model long-distance dependencies while ensuring controllable computational cost. The structural principle diagram of the convolution-transformer backbone network is shown below. Figure 2 As shown.

[0032] First, the input sequence obtained in step S1 is... By cascading several levels of one-dimensional convolution and pooling, the temporal domain is downsampled and channels are expanded, resulting in a sequence representation with significantly shortened length: (5) in This represents the sequence operator "convolution → nonlinearity → pooling". Indicates length, This represents the number of channels. After downsampling, each row yields a token, and each token aggregates local patterns from adjacent time slices; the convolution stride and pooling scale jointly determine this. The compression ratio is adjusted to reduce the complexity of subsequent attention while preserving the main energy bands.

[0033] Then, a lightweight Transformer encoder is used to... To perform global interaction modeling, the input is reduced to a lightweight Transformer encoder, which consists of a multi-head self-attention (MHA), residual and layer normalization (LN), feedforward network (FFN), and residual and layer normalization (LN).

[0034] Bullish Self-Attention Query, key, and value can be obtained through learnable linear mapping: (6) Among them, by number of heads Divided, each head has a dimension of For the first Each attention point, calculate: (7) The outputs from each head are concatenated by channel and linearly mapped back. Wei Dedao Then, random loss is applied, residual stacking is performed, and layer normalization is applied: (8) This step establishes a long-range dependency between any two tokens, which increases the computational cost, since the preceding convolution has already... Compress to Therefore, the overall computation is controllable. No explicit positional encoding is introduced in the implementation; the temporal order is implicitly preserved by convolutional downsampling and attention weights.

[0035] In feedforward networks (FFNs) Apply two fully connected and non-linear layers to each token, first increasing the dimensionality and then restoring it to its original state. Dimensions, with random inactivation in the middle to suppress overfitting: (9) (10) in It is ReLU. Then, residual stacking and layer normalization are performed to obtain the final encoded representation: (11) This approach preserves the local robustness of convolution while establishing semantic alignment across the entire time domain through attention, ensuring stability and feature discriminativeness during deep training.

[0036] To further consolidate discriminative clues and reduce the size of classification head parameters, in the encoding representation Then, a one-dimensional convolutional layer is introduced to perform learnable temporal convergence: (12) Then Flattened into vectors, high-order discriminative features are extracted using two layers of fully connected units with ReLU, outputting the occupancy rate and class probability: (13) in It consists of several fully connected and nonlinear units. During the training phase, cross-entropy is used for end-to-end supervised optimization. (14) in For the number of categories, To provide an effective supervisory label, the optimizer employs an adaptive gradient method.

[0037] Step S3 specifically includes: adding a temporal decoding module to the one-dimensional convolution and Transformer backbone network in step S2 to form a one-dimensional masked autoencoder (MAE), and performing self-supervised pre-training on the unlabeled DAS long-term sequence. The overall process is shown in Figure 3(a), and the loss function calculation is shown in Figure 3(b).

[0038] First, the normalized single-channel DAS sequence The input to the convolutional downsampling front-end (stem) of the backbone network in step S2 compresses the long sequence into several small token segments of equal length in the time dimension according to a pre-defined sliding and pooling method. This transformation locally condenses the original long sequence of forms into a few key parts, while preserving the original sense of order for subsequent global relationship modeling.

[0039] At the token level, randomly select a subset of locations for "occlusion". Let the set of occluded locations be denoted as . (Its size is approximately the preset ratio) For the occluded positions, replace the original token with a placeholder vector; leave the other positions unchanged, resulting in a sequence containing only visible information. This is equivalent to "information erasure" at the time segment level, forcing the model to rely on the near and far context to infer the erased parts.

[0040] Next, Input the same Transformer encoder as in step S2 Output hidden representation The encoding end establishes long-range dependencies among the remaining few visible tokens through multi-head self-attention. It comprehensively considers similar waveforms generated when the vehicle passes through various sampling locations, energy transfer and phase correlation across tokens, and statistical consistency of background noise to fill in the occluded parts. Compared with simple convolutional imputation, it does not require a fixed receptive field of view in advance. It can automatically fuse evidence for different durations and different working states, thereby improving the capture effect of long-term dependencies and cross-segment correlation phenomena.

[0041] To restore the encoded global representation back to the token-level signal, a lightweight decoder is set up. Perform a token-by-token nonlinear mapping (two fully connected layers: ReLU followed by linear mapping) at each time point to obtain the reconstructed sequence. The training objective is to measure the reconstruction error only at the masked locations, specifically using the mask mean square error: (15) in The first part is divided into reality and reconstruction. One token, To prevent extremely small constants with a denominator of zero, an adaptive gradient method is used during optimization. The parameters are jointly minimized, and convergence can be monitored in conjunction with commonly used adaptive optimizers and validation sets.

[0042] Through the "occlusion-reconstruction" task, the encoder is forced to learn general temporal characteristics such as vehicle dynamics, media propagation, and occupancy changes from the statistical structure of visible segments. Once the pre-training is complete, the decoder is removed, but the backbone parameters are retained and initialized. Then, a classification head is added to a small amount of labeled data for supervised fine-tuning, thereby improving the initial performance and convergence stability of occupancy recognition without affecting the backbone structure.

[0043] Under this training model, the model must infer the temporal trend of the occluded segment from the statistical structure of the visible segment: when a vehicle passes by, the relevant waveforms appearing along the fiber sampling points will exhibit transferable local-global patterns across multiple tokens; when external noise or non-target disturbances occur, their inconsistency at different locations prompts the encoder to learn a more discriminative representation. Thus, the encoder acquires several general time-series characteristics regarding vehicle dynamics, medium propagation paths, and changes in vehicle interior occupancy from a large amount of unlabeled DAS data. After pre-training, the decoder is removed, and the remaining data is retained. and Then, supervised fine-tuning is performed on a small amount of labeled data, which can significantly improve the initial accuracy and convergence stability of occupancy classification without changing the main framework.

[0044] Step S4 specifically includes: after completing the self-supervised pre-training in step S3, removing the decoding end of MAE, leaving only the encoding backbone composed of one-dimensional convolution and Transformer as the feature extraction network, and adding a classification head for vehicle occupancy recognition at its output position.

[0045] The backbone network trained in step S3 now possesses the ability to perform joint "local-global" modeling of long-term temporal DAS sequences. The labeled samples obtained in step S1 are now... By directly inputting the backbone, compact temporal features are obtained after convolutional downsampling and self-attention processing. : (16) in This represents the downsampling front-end for convolution and pooling. This represents the Transformer encoder. To balance stability and adaptability, the backbone parameters can be fully fine-tuned or partially frozen and the learning rate adjusted in stages. This approach ensures that the pre-trained representations are not corrupted while allowing them to be appropriately adapted to labeled tasks.

[0046] exist Then, a one-dimensional convolutional or fully connected compression layer is applied to linearly combine and locally aggregate adjacent time slices, resulting in a shorter sequence or a single converged vector. Compared to simple pooling, this compression layer strikes a balance between preserving fine-grained transient information and computational overhead. After flattening, the vector is passed sequentially through several fully connected layers with non-linear activation to obtain the category scoring vector. Then, Softmax normalization was used to obtain the predicted probability for each passenger number category: (17) To mitigate the risks of overconfidence and overfitting associated with rigid one-hot labels, label smoothing is introduced. This smoothing is applied to the true categories. The label was softened from 1 to ,the remaining The category was upgraded from 0 to Result in softened label Training uses cross-entropy based on smooth labels: (18) This operation not only suppresses overconfidence in a single category, but also reduces the amplification effect of occasional label noise on the gradient, and enhances the robustness of identification in relatively rare conditions such as multi-occupant situations.

[0047] Then, the labeled DAS samples constructed in step S1 are fed into the model using the same channel selection, time window, and normalization methods as in the pre-training stage. Probabilities are obtained through forward propagation in mini-batch mode. The loss function is calculated using the above formula and backpropagated to jointly update the backbone and classifier head parameters. Classification accuracy and other metrics are monitored synchronously on an independent validation set, and an early stopping strategy is used to control the number of training epochs: when the validation metrics no longer improve within a certain number of epochs, training is stopped and the model is rolled back to the optimal weights. This process effectively avoids overfitting and ensures the model remains robust under different road conditions, vehicle speeds, and passenger loads.

[0048] Step S5 specifically includes: In the model deployment stage, the distributed fiber optic sensing device is placed in a specific required location. The real-time collected DAS time series are numerous and complex. After preprocessing in step S1, data that can be input into the model is obtained.

[0049] Each time window Input the "convolutional + Transformer" backbone network trained in step S4 to extract compact temporal representations. Subsequently, the classification head is mapped to the scoring vector of each passenger number category. The probability distribution is obtained through Softmax: (19) in The total number of categories, This indicates that the time window belongs to the first... Confidence level of the number of passengers in each class. Deployment end output. As the result of identifying the occupancy rate of this window, while retaining As a confidence index, it facilitates subsequent fusion and threshold control.

[0050] When the same vehicle passes through, the system typically obtains several temporally consecutive and spatially contiguous window predictions. By combining the continuity of vehicle passage time and trajectory, vehicles belonging to the same target can be identified. Probabilistic fusion is performed across multiple windows to obtain a more robust single-pass estimation. Weighted average fusion is employed. (20) Among them, weight can be The proximity of the window to the vehicle center at that moment, or the signal-to-noise ratio, is determined. The final output category is... This fusion method, without altering the model structure, utilizes the spatiotemporal continuity of vehicle trajectories to suppress occasional noise in a single window, thereby improving the discrimination stability under multi-occupant and weak signal conditions.

[0051] To further improve system reliability, a minimum confidence threshold can be set at the deployment end. .when If this occurs, mark the estimate as "low confidence" and handle the anomaly accordingly.

[0052] Since the entire process relies solely on DAS timing signals without acquiring or storing images, it can achieve non-contact, privacy-friendly intelligent detection of vehicle occupancy rates while meeting real-time requirements.

[0053] It will be understood by those skilled in the art that the above descriptions are merely preferred examples of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A method for estimating the number of vehicle occupants based on DAS time-series signals, comprising the following steps: S1: Use a distributed fiber optic acoustic sensing system laid under the road to obtain the original one-dimensional vibration signal generated when a vehicle passes by. Perform operations such as channel selection, window truncation and marking on the original fiber optic data to form a one-dimensional time series sample library with vehicle occupancy indicators. S2: Construct a feature extraction backbone network combining one-dimensional convolution and Transformer to perform hierarchical downsampling and feature extraction on long-sequence one-dimensional DAS signals, transforming the original long-sequence vibration signal into several shorter feature token sequences, providing a consistent backbone feature space for subsequent MAE pre-training and supervised fine-tuning. S3: Without labels, the backbone network constructed in step S2 is used as an encoder. A decoding module is added to form a complete masked autoencoder (MAE). The MAE decoder is used to perform a sequence reconstruction task after random masking to mine general temporal features in a large amount of unlabeled DAS data and optimize the MAE encoder parameters. Self-supervised pre-training is performed on the DAS time series samples preprocessed in S1. The general temporal feature representation of vehicle passage is learned and obtained by relying on random masking and reconstruction tasks. S4: Based on the DAS dataset with labels of the number of occupants in the vehicle, the MAE decoder is removed, and only the pre-trained one-dimensional convolutional + Transformer backbone network (i.e. the MAE encoder) is retained. A classification head is added to the end of its output, and supervised fine-tuning training is performed using label smoothing to obtain a multi-classification model specifically for recognizing vehicle occupancy. S5: Use the newly acquired DAS time series as input to the vehicle occupancy classification model trained in step S4, and output the number of occupants in the vehicle for each time window to achieve the purpose of automatic identification of the vehicle occupancy rate of passing vehicles.

2. A method for estimating the number of vehicle occupants based on DAS time-series signals, characterized in that: Step S1 specifically includes: using a distributed optical fiber sensing system laid under the road to collect multi-channel one-dimensional vibration signals of the target road segment during vehicle passage, and saving the collected raw data in the form of a time-series array; selecting several continuous spatial channels covering the vehicle trajectory as the range of interest channels based on the known vehicle lane position and driving direction, and windowing the raw DAS signal on the time axis according to the start and end times of the vehicle passage event to obtain a fixed-length one-dimensional time series segment; further, combining the vehicle number, driving speed, driving direction and number of occupants recorded in the experimental conditions, pairing each time series segment with the vehicle, and classifying the samples into the corresponding occupancy rate category according to the number of occupants, thereby constructing a training sample set with the DAS one-dimensional time series as input and the occupant number category as the output label; performing zero-padding or padding on the sequence segments with insufficient time length to ensure that all samples maintain a consistent length in the time dimension to adapt to the input requirements of subsequent neural network training.

3. A method for estimating the number of vehicle occupants based on DAS time-series signals, characterized in that: Step S2 specifically includes: based on the labeled DAS one-dimensional time series obtained in step S1, constructing a feature extraction backbone network consisting of a multi-layer one-dimensional convolutional network and a Transformer encoder connected in series; firstly, using several levels of one-dimensional convolutional layers in conjunction with pooling or strided convolution operations, performing hierarchical downsampling on the original long sequence. Each one-dimensional convolutional layer extracts local vibration patterns in the time direction, and gradually shortens the sequence length and increases the number of channels, thereby suppressing high-frequency noise while preserving the key signal of the vehicle passing through. After multiple downsampling steps, the original length is reduced to... The DAS sequence is compressed to a length of The intermediate feature sequence is used, where the feature vectors at each time point are treated as a set of low-dimensional feature tokens. Then, this intermediate feature sequence is input into the Transformer encoder, which has a length of [missing information]. Multi-head self-attention operations and feedforward network transformation operations are performed on the token sequence to model the long-distance dependencies between different time positions and the overall energy distribution characteristics. By utilizing the self-attention mechanism, the model can simultaneously capture the comprehensive features of key time periods such as vehicle entry, passage, and departure from the sensing area, and obtain feature token representations containing overall vehicle dynamic information and occupancy-related patterns, thus providing a consistent backbone feature space for subsequent self-supervised pre-training and supervised classification.

4. A method for estimating the number of vehicle occupants based on DAS time-series signals, characterized in that: Step S3 specifically includes: Based on the one-dimensional convolutional and Transformer backbone network described in step S2, a time-scale decoding module is added, which, together with the backbone network, constitutes a one-dimensional masked autoencoder (MAE) structure for self-supervised pre-training of unlabeled DAS time series. First, each DAS time series is cut into continuous time patches of a fixed length. Then, according to a preset masking ratio, most patches are randomly selected for masking, with only a small number of patches retained as visible tokens. The masked patches are replaced with special mask markers. Next, the masked time series is input into the one-dimensional convolutional + Transformer encoder of step S2 to obtain a compressed latent space feature representation. Then, the decoding end performs upsampling operations layer by layer to restore the time dimension, thereby reconstructing the complete DAS time series segment. Finally, the parameters of the masked autoencoder are optimized and adjusted using the mean squared error between the reconstructed sequence and the original unmasked sequence as the reconstruction loss function.

5. A method for estimating the number of vehicle occupants based on DAS time-series signals, characterized in that: Step S4 specifically includes: after completing the self-supervised pre-training described in step S3, removing the masked autoencoder and retaining the encoding backbone network composed of one-dimensional convolution and Transformer as the feature extractor; adding a classification head for identifying vehicle occupancy rate at the end of the backbone network, which includes: a one-dimensional convolutional or fully connected compression layer for further integrating time-dimensional features and compressing them to a smaller scale, followed by several fully connected layers and non-linear activation functions, and finally using Softmax to output the predicted probability of each occupant number category; in the supervised fine-tuning stage, using the DAS time series with occupant number labels constructed in step S1... Using samples as input, the model is trained using the cross-entropy loss function. A label smoothing strategy is also incorporated to soften the original one-hot labels, suppressing overconfidence in a single class and reducing the impact of individual sample label noise on the training process. This improves the macro-average F1 score and enhances the robustness of recognition in low-occupancy conditions such as multi-occupant scenarios. During training, comprehensive indicators such as macro-average F1 and classification accuracy are monitored on the validation set, and an early stopping strategy is used to select appropriate training rounds. This results in a vehicle occupancy classification model suitable for various application scenarios and different occupant numbers.

6. A low-overlap point cloud registration method based on deep learning, characterized in that: Step S5 specifically includes: In the model deployment phase, the new DAS time series acquired online by the distributed fiber optic acoustic wave sensing system is processed according to the same preprocessing procedure as the training phase, including channel selection, time window truncation, and standardization, and then input into the vehicle occupancy classification model trained in step S4; The model first extracts temporal features through one-dimensional convolution and Transformer backbone network, and then outputs the predicted probability of each occupant number category by the classification head, and determines the category with the highest probability as the vehicle occupancy rate corresponding to that time window; Based on the vehicle passage time and trajectory information, the prediction results of multiple windows can be fused to obtain the final occupant number estimate for a single vehicle passage, realizing real-time, non-contact, and image-free intelligent detection of the vehicle occupancy rate of passing vehicles.