Subway positioning method, device, system, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610673003.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-25
AI Technical Summary
现有技术中,在地铁车厢内对移动终端进行连续定位的方案主要存在以下缺陷:其一,卫星-惯性混合方案,由于隧道内卫星信号完全被遮挡,定位信息会出现物理层中断,且普通移动终端内置的低精度惯性器件在长时间无校准的情况下会产生严重的累积误差,无法实现长区间的精准定位;其二,移动蜂窝-指纹方案或蓝牙信标定位方案,不仅需要大规模改造地铁基础设施,建设与维护成本极高,而且射频信号在狭长的金属隧道环境及列车金属车体内会发生严重的法拉第屏蔽和多径反射效应,导致信号强度波动剧烈、多普勒频移显著,使得定位结果极不稳定、跳变频繁,实时性与空间区分度均较差
[0022]第八方面,本申请实施例提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现第一方面所述的地铁定位方法,或实现第二方面所述的地铁定位方法。
Smart Images

Figure CN122815328A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mobile terminal positioning in urban rail transit environments, and more particularly to a subway positioning method, device, system, electronic device, and storage medium. Background Technology
[0002] With the rapid development of urban rail transit, subways have become the main commuting tool for many students. However, the subway network is complex and intricate. To ensure students' travel safety and prevent abnormal situations such as boarding the wrong train, going in the wrong direction, or missing their stop, it is necessary to continuously and accurately locate the mobile devices carried by students inside the train carriages so that parents or schools can be alerted in a timely manner if the route deviates.
[0003] However, the subway environment (especially inside tunnels) is a typical enclosed and confined space. Existing technologies for continuous positioning of mobile terminals inside subway cars suffer from the following drawbacks: First, the satellite-inertial hybrid approach suffers from physical layer interruptions in positioning information due to complete satellite signal blockage within tunnels. Furthermore, the low-precision inertial devices built into ordinary mobile terminals accumulate significant errors over extended periods without calibration, making accurate positioning over long distances impossible. Second, mobile cellular-fingerprint or Bluetooth beacon positioning solutions require large-scale modifications to subway infrastructure, resulting in extremely high construction and maintenance costs. Moreover, radio frequency signals experience severe Faraday shielding and multipath reflection effects in the narrow metal tunnel environment and within the metal body of the train, leading to drastic signal strength fluctuations and significant Doppler shifts. This results in highly unstable and frequently fluctuating positioning results, with poor real-time performance and spatial resolution.
[0004] Therefore, how to achieve continuous, stable, and high-precision positioning in a closed subway environment without increasing the cost of additional subway infrastructure construction is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a subway positioning method, device, system, electronic device, and storage medium, which enables continuous, stable, and high-precision positioning in a closed subway environment without increasing additional subway infrastructure construction costs.
[0006] In a first aspect, embodiments of this application provide a subway positioning method, applied to a server, comprising: Obtain the target acoustic feature vector of the subway environment audio where the target object is located; The target acoustic feature vector is matched with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic signature database is constructed based on the first subway feature noise collected in the subway tunnel, and the first subway feature noise includes at least one of wheel-rail noise and wind noise; Based on the candidate location set, the target location of the target object is determined.
[0007] In one embodiment, before matching the target acoustic feature vector with a subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object, the method further includes: Acquire sample audio and corresponding geographical location information from multiple collection points within a subway tunnel, and acquire interference noise samples; the sample audio includes wheel-rail noise and / or wind noise, and the interference noise samples include at least one of human voice, subway broadcasts, door opening and closing sounds, and construction noise; Based on the interference noise samples, the sample audio is processed, and the sample acoustic feature vector is extracted from the processed sample audio. The subway line is divided into multiple geographic grids, and the sample acoustic feature vectors are assigned to the corresponding geographic grids based on the geographic location information. The acoustic signature database for the subway tunnel environment is obtained by training based on the segmented sample acoustic feature vectors.
[0008] In one embodiment, processing the sample audio based on the interference noise sample and extracting the sample acoustic feature vector from the processed sample audio includes: Based on the interference noise samples, the sample audio is sequentially subjected to noise reduction and normalization processing to obtain standard audio; The standard audio is segmented into frames to obtain a standard audio frame sequence; Feature extraction is performed on the standard audio frame sequence to obtain a time-frequency acoustic feature sequence; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum feature; Based on the time-frequency acoustic feature sequence, the sample acoustic feature vector is generated by aggregation.
[0009] In one embodiment, the step of performing noise reduction and normalization processing on the sample audio based on the interference noise samples to obtain standard audio includes: The sample audio is converted to a time-frequency signal to obtain a sample time-frequency spectrum. Based on the interference noise samples, identify the first random interference noise in the sample audio and identify the first subway characteristic noise in the sample audio; The first random interference noise is filtered, and the target frequency band of the first subway characteristic noise is enhanced to extract the target subway characteristic parameters. The target subway feature parameters and the sample time-frequency graph are input into a preset neural network model to obtain the feature mask output by the preset neural network model. Based on the feature mask, the time-frequency spectrum of the sample is weighted and reconstructed and transformed in the time domain to obtain the processed sample audio. The processed sample audio is then subjected to global level normalization and spectral tilt compensation in sequence to obtain the standard audio.
[0010] In one embodiment, the step of training based on the segmented sample acoustic feature vectors to obtain the subway tunnel environmental acoustic signature database includes: The acoustic feature vectors of samples from multiple geographic grids belonging to the same metro line are aggregated and trained to generate a line-level acoustic model corresponding to the metro line. The acoustic feature vectors of samples from multiple geographic grids belonging to the same driving section are aggregated and trained to generate a section-level acoustic model corresponding to the driving section. The grid-level acoustic model corresponding to the geographic grid is generated by training based on the acoustic feature vectors of samples divided into the same geographic grid. The subway tunnel environmental acoustic model database includes the line-level acoustic model, the section-level acoustic model, and the grid-level acoustic model.
[0011] In one embodiment, the subway tunnel environment acoustic signature database includes a line-level acoustic model, a section-level acoustic model, and a grid-level acoustic model. Matching the target acoustic feature vector with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object includes: The target acoustic feature vector is matched with the line-level acoustic model to determine the candidate line set; The target acoustic feature vector is matched with the interval-level acoustic model corresponding to the candidate line set to determine the candidate interval set; The acoustic feature vectors are matched with the grid-level acoustic models corresponding to the candidate interval set to determine the candidate location set.
[0012] In one embodiment, determining the target location of the target object based on the candidate location set includes: Obtain the historical location and first timestamp of the target object; Based on the second timestamp corresponding to the target acoustic feature vector, the historical positioning location, and the preset maximum running speed of the first timestamp, the theoretically achievable range is determined. Based on the theoretically reachable range, the candidate location set is filtered to obtain the remaining candidate locations; Based on the remaining candidate locations, the target location is determined.
[0013] In one embodiment, determining the target location based on the remaining candidate locations includes: Obtain the operational constraint information of the subway environment in which the target object is located; the operational constraint information includes at least one of path topology, speed limit rules, and direction of travel; Based on the operational constraint information, the voiceprint matching scores corresponding to the remaining candidate positions are weighted to obtain a comprehensive confidence score. Based on the comprehensive confidence score, the target location is determined from the remaining candidate locations.
[0014] Secondly, embodiments of this application provide a subway positioning method applied to a mobile terminal, comprising: Obtain the target acoustic feature vector of the subway environment audio at the target object location; Send the target acoustic feature vector to the server; The server receives the target location returned by the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database; wherein the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
[0015] In one embodiment, obtaining the target acoustic feature vector of the subway environment audio where the target object is located includes: Collect audio from the current subway environment; Identify the second random interference noise, fixed equipment noise, and second subway characteristic noise in the current subway environment audio; The second random interference noise and the fixed equipment noise are filtered, and the target frequency band of the second subway characteristic noise is enhanced to obtain the target audio. Based on the target audio, the target acoustic feature vector is extracted.
[0016] In one embodiment, obtaining the target acoustic feature vector of the subway environment audio where the target object is located includes: In the first working mode, target sound events in the subway environment where the target object is located are monitored; When the target sound event is detected, the system switches from the first working mode to the second working mode. In the second working mode, the current subway environment audio is collected, and the target acoustic feature vector of the current subway environment audio is extracted.
[0017] Thirdly, embodiments of this application provide a subway positioning device, comprising: The first acquisition module is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located; The feature matching module is used to match the target acoustic feature vector with the subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise; The location determination module is used to determine the target location of the target object based on the candidate location set.
[0018] Fourthly, embodiments of this application provide a subway positioning device, comprising: The second acquisition module is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located; The feature sending module is used to send the target acoustic feature vector to the server; The location receiving module is used to receive the target location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
[0019] Fifthly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the subway positioning method described in the first aspect, or the subway positioning method described in the second aspect.
[0020] Sixthly, embodiments of this application provide a subway positioning system, characterized in that it includes a server and a mobile terminal, wherein: The server is used to execute the subway positioning method as described in the first aspect; The mobile terminal is used to perform the subway positioning method as described in the second aspect.
[0021] In a seventh aspect, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the subway positioning method described in the first aspect, or implements the subway positioning method described in the second aspect.
[0022] Eighthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the subway positioning method described in the first aspect, or implements the subway positioning method described in the second aspect.
[0023] The subway positioning method, device, system, electronic device, and storage medium provided in this application obtain the target acoustic feature vector of the subway environment audio where the target object is located through a server; then, the target acoustic feature vector is matched with a subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, the first subway characteristic noise including at least one of wheel-rail noise and wind noise; and then, based on the candidate location set, the target positioning location of the target object is determined. This application embodiment uses the inherent characteristic noise in the subway tunnel environment as an acoustic fingerprint for positioning, without relying on external satellite signals, overcoming the problem of physical layer interruption of positioning caused by satellite signal blockage in the tunnel; at the same time, this method also does not require large-scale deployment of external hardware facilities such as Bluetooth beacons or cellular base stations in the subway tunnel, thereby avoiding high construction and maintenance costs. In addition, compared with radio frequency signals, acoustic features are minimally affected by Faraday shielding of the metal car body and tunnel multipath reflection effects, avoiding the problems of drastic signal strength fluctuations and frequent jumps in positioning results in traditional solutions, thus improving the positioning accuracy and stability in the closed subway environment. In summary, the embodiments of this application achieve continuous, stable, and high-precision positioning in a closed subway environment without increasing the cost of subway infrastructure construction. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is one of the flowcharts illustrating the subway positioning method provided in the embodiments of this application.
[0026] Figure 2 This is the second flowchart illustrating the subway positioning method provided in the embodiments of this application.
[0027] Figure 3 This is the third flowchart illustrating the subway positioning method provided in the embodiments of this application.
[0028] Figure 4 This is the fourth flowchart of the subway positioning method provided in the embodiments of this application.
[0029] Figure 5 This is the fifth flowchart illustrating the subway positioning method provided in the embodiments of this application.
[0030] Figure 6 This is the sixth flowchart illustrating the subway positioning method provided in the embodiments of this application.
[0031] Figure 7 This is one of the structural schematic diagrams of the subway positioning device provided in the embodiments of this application.
[0032] Figure 8 This is the second structural schematic diagram of the subway positioning device provided in the embodiments of this application.
[0033] Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] With the rapid development of urban rail transit, subways have become the main commuting tool for many students. However, the subway network is complex and intricate. To ensure students' travel safety and prevent abnormal situations such as boarding the wrong train, going in the wrong direction, or missing their stop, it is necessary to continuously and accurately locate the mobile devices carried by students inside the train carriages so that parents or schools can be alerted in a timely manner if the route deviates.
[0036] Currently, existing technologies for continuous positioning of mobile terminals in subway cars mainly include satellite-inertial hybrid schemes, mobile cellular-fingerprint schemes, and Bluetooth beacon positioning schemes.
[0037] The satellite-inertial hybrid scheme refers to a scheme that primarily uses GNSS (Global Navigation Satellite System) and secondarily uses MEMS (Micro-Electro-Mechanical System) inertial navigation. Specifically, when members enter a station and during train operation, GNSS is used to reset inertial drift during the occasional moments when the subway exits the tunnel. After entering the tunnel, inertial integration is used to maintain the position. This scheme has the following drawbacks: First, if the satellite signal is zero inside the tunnel, the positioning information will experience a physical layer interruption, and the terminal device cannot obtain independent position updates at any time. Second, the zero-bias drift of the inertial device causes the error to accumulate exponentially over time, and the terminal-level MEMS alone cannot maintain an accuracy of <100m over long intervals (>3min). In addition, it is necessary to rely on the train end or a high-level inertial navigation system for global correction. Ordinary terminals have neither a high-precision IMU (Inertial Measurement Unit) nor an external correction source, making resource accessibility zero. In other words, this solution relies entirely on available satellite time slots and high-precision inertial navigation devices, and is not applicable to terminal equipment without satellites or inertial navigation.
[0038] The mobile cellular fingerprinting scheme involves deploying leaky cables, small base stations, or repeaters within tunnels to associate RSSI / TOA (Received Signal Strength Indicator / Time of Arrival) fingerprints within the train carriages with the line mileage, forming a tunnel fingerprint database. Terminals can obtain relative mileage by measuring base station signals in real time and matching the fingerprints. This scheme has the following drawbacks: First, the base station density within tunnels is low, with a single-point coverage radius >300m, resulting in insufficient spatial discrimination and excessively large GDOP (Geometric Dilution of Precision), often leading to positioning errors >150m. Second, the train's metal body creates Faraday shielding, causing RSSI differences >15dB between inside and outside the carriage, resulting in ambiguous matches for the same fingerprint at different door locations. Furthermore, the combined effects of multipath and Doppler effects cause the instantaneous RSSI standard deviation >8dB, requiring continuous sampling >30s for convergence, resulting in poor real-time performance.
[0039] The Bluetooth beacon positioning scheme involves deploying a low-power Bluetooth beacon every 5-10 meters along the wall of the tunnel or platform, continuously broadcasting the UUID (Universally Unique Identifier), Major, Minor, and RSSI. After scanning, the terminal device uses the RSSI-distance model or triangulation / fingerprint algorithm to calculate the location. The proposed solution has the following drawbacks: First, it requires large-scale modifications to the subway infrastructure, including the deployment of beacons along the entire line and regular battery replacements, significantly increasing construction and maintenance costs. Second, the 2.4GHz signal experiences multiple reflections between the tunnel walls and the train body, resulting in RSSI fluctuations of σ≈10dB, leading to distance estimation errors exceeding 50% and frequent changes in positioning results. Third, the beacon coverage radius is only 10-30m, requiring the deployment of over 1000 nodes over long sections, complicating equipment management. Furthermore, each node broadcasts periodically, causing the air channel occupancy rate to increase linearly with the number of nodes, resulting in co-channel interference and extremely high packet loss rates. Additionally, under the battery-powered model, the Bluetooth beacon transmission current exceeds 8mA, necessitating batch replacements every 1-2 years. If cable power is used, the low-power advantage is lost, and maintenance energy costs are converted into system-level power consumption.
[0040] Therefore, how to achieve continuous, stable, and high-precision positioning in a closed subway environment without increasing the cost of additional subway infrastructure construction is a technical problem that urgently needs to be solved by those skilled in the art.
[0041] This application proposes a subway positioning method, device, system, electronic device, and storage medium, which are described below in conjunction with... Figures 1-9 Describe it.
[0042] Figure 1 This is one of the flowcharts illustrating the subway positioning method provided in the embodiments of this application, such as... Figure 1 As shown, the subway positioning method includes steps S110, S120 and S130.
[0043] Step S110: Obtain the target acoustic feature vector of the subway environment audio where the target object is located.
[0044] The subway positioning method provided in this application is applied to the server.
[0045] The subway environment includes, but is not limited to: subway cars, subway platforms, and subway transfer passages.
[0046] The audio of the subway environment where the target is located can be collected by the target's mobile device. The mobile device can include, but is not limited to: learning Pads (tablets), smart student cards, smartwatches, smartphones, tablets, etc.
[0047] A target acoustic feature vector refers to the transformation of physical environment audio into a low-dimensional, compact, and highly discriminative mathematical representation. In one embodiment, the acoustic feature vector of the subway environment audio in which the target object is located is extracted by a mobile terminal and denoted as the target acoustic feature vector. This target acoustic feature vector is then sent to the server via a mobile network. That is, the mobile terminal does not transmit the original audio that may involve user privacy over the mobile network; instead, it only uploads the de-identified and highly abstracted mathematical feature vector.
[0048] Furthermore, when sending the target acoustic feature vector, the mobile terminal can package and send information such as the mobile terminal's device ID (Identifier) and current timestamp together to the server.
[0049] Correspondingly, the server receives the target acoustic feature vector of the subway environment audio of the target object sent by the mobile terminal.
[0050] Step S120: Match the target acoustic feature vector with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic signature database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
[0051] After obtaining the voiceprint feature vector, the server performs a rapid search in a pre-established subway tunnel environment voiceprint database to match the target acoustic feature vector with the subway tunnel environment voiceprint database, thereby determining the candidate location set corresponding to the target object.
[0052] The subway tunnel environmental acoustic signature database is a geolocation-acoustic model mapping database built based on subway characteristic noise with strong location specificity. To distinguish it from the subway characteristic noise during the extraction of the target acoustic feature vector, the subway characteristic noise collected within the subway tunnel is designated as the first subway characteristic noise. The first subway characteristic noise includes at least one of wheel-rail noise and wind noise. Wheel-rail noise refers to the friction / impact sound generated by the specific wear state of the wheels and rails, while wind noise refers to the aerodynamic wind noise formed by the influence of tunnel diameter and wall roughness.
[0053] In one embodiment, the construction process of the subway tunnel environmental acoustic signature database is as follows: Sample audio data and corresponding geographical location information from multiple collection points within the subway tunnel are acquired, along with interference noise samples. The sample audio data includes wheel-rail noise and / or wind noise, while the interference noise samples include at least one of human voice, subway announcements, door opening / closing sounds, and construction noise. Based on the interference noise samples, the sample audio data is processed, and sample acoustic feature vectors are extracted from the processed sample audio data. The subway line is divided into multiple geographic grids, and the sample acoustic feature vectors are assigned to the corresponding geographic grids based on the geographical location information. Training is performed based on the sample acoustic feature vectors assigned to the same geographic grid to generate a grid-level acoustic model corresponding to the geographic grid, thus forming the subway tunnel environmental acoustic signature database.
[0054] In another embodiment, the construction process of the subway tunnel environmental acoustic vector database is as follows: Sample audio data and corresponding geographical location information from multiple collection points within the subway tunnel are acquired, along with interference noise samples. The sample audio data includes wheel-rail noise and / or wind noise, while the interference noise samples include at least one of human voice, subway broadcasts, door opening / closing sounds, and construction noise. Based on the interference noise samples, the sample audio data is processed, and sample acoustic feature vectors are extracted from the processed sample audio data. The subway line is divided into multiple geographic grids, and the sample acoustic feature vectors are assigned to the corresponding geographic grids based on the geographical location information. Training is performed based on the sample acoustic feature vectors assigned to the same geographic grid to generate a grid-level acoustic model corresponding to the geographic grid. The sample acoustic feature vectors from multiple geographic grids belonging to the same operating section are aggregated and trained to generate an interval-level acoustic model corresponding to the operating section. The sample acoustic feature vectors from multiple geographic grids belonging to the same subway line are aggregated and trained to generate a line-level acoustic model corresponding to the subway line. The subway tunnel environmental acoustic vector database includes line-level acoustic models, interval-level acoustic models, and grid-level acoustic models.
[0055] The specific construction process can be found in the following examples, which will not be elaborated here.
[0056] Furthermore, during the matching process, one or more candidate locations that meet preset conditions can be selected by calculating the voiceprint matching score, forming a candidate location set. The voiceprint matching score can be represented by similarity.
[0057] Step S130: Based on the candidate location set, determine the target location of the target object.
[0058] Finally, the server determines the target location of the target object based on the candidate location set.
[0059] In one embodiment, the candidate position with the highest voiceprint matching score in the candidate position set is directly determined as the target location of the target object.
[0060] In another implementation, intelligent judgment is made based on the candidate location set and other clues. This includes, but is not limited to: (1) temporal context filtering: referring to the historical location of the mobile terminal device, candidate locations that cannot be reached within a reasonable time are eliminated; (2) operation logic verification and weighting: using information such as subway route maps, timetables, and operating directions, the candidate locations are weighted according to their reasonableness. For example, the subway management system is accessed, and impossible vehicle location coordinates are eliminated by using real-time vehicle location information obtained from it. By integrating multiple information such as voiceprint, time, and operation logic, the server will finally calculate the candidate location with the highest comprehensive confidence score and determine it as the target location of the target object.
[0061] Furthermore, once the server determines the target location of the object, it sends the target location to the corresponding mobile terminal. The representation of the target location can include, but is not limited to, latitude and longitude, and the route / section / station where it is located. In addition, while sending the target location, it can also push the target location along with the mobile terminal's device ID, voiceprint matching score, and overall confidence score to relevant monitoring or alarm systems, such as parental monitoring apps and educational early warning systems, so that alarms can be triggered promptly in case of abnormal situations.
[0062] Of course, it should be understood that authorization from the relevant user should be obtained before pushing the message to the relevant monitoring or alarm system.
[0063] Furthermore, the device ID of the mobile terminal can be associated with and saved with the target location, current timestamp, voiceprint matching score, and overall confidence score for later retrieval.
[0064] The subway positioning method provided in this application obtains the target acoustic feature vector of the subway environment audio where the target object is located through a server; then, it matches the target acoustic feature vector with a subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, the first subway characteristic noise including at least one of wheel-rail noise and wind noise; and then, based on the candidate location set, the target positioning location of the target object is determined. This application embodiment uses the inherent characteristic noise in the subway tunnel environment as an acoustic fingerprint for positioning, without relying on external satellite signals, overcoming the problem of physical layer interruption in positioning caused by satellite signal blockage in the tunnel; at the same time, this method also does not require large-scale deployment of external hardware facilities such as Bluetooth beacons or cellular base stations in the subway tunnel, thereby avoiding high construction and maintenance costs. In addition, compared with radio frequency signals, acoustic features are minimally affected by Faraday shielding of the metal car body and tunnel multipath reflection effects, avoiding the problems of drastic signal strength fluctuations and frequent jumps in positioning results in traditional schemes, thus improving the positioning accuracy and stability in the closed subway environment. In summary, the embodiments of this application achieve continuous, stable, and high-precision positioning in a closed subway environment without increasing the cost of subway infrastructure construction.
[0065] Based on any of the above embodiments Figure 2 This is a second flowchart illustrating the subway positioning method provided in this application embodiment, as shown below. Figure 2 As shown, before step S120, the method further includes steps S10, S20, S30 and S40.
[0066] Step S10: Obtain sample audio and corresponding geographical location information from multiple collection points within the subway tunnel, and obtain interference noise samples; the sample audio includes wheel-rail noise and / or wind noise, and the interference noise samples include at least one of human voice, subway broadcast, door opening and closing sounds, and construction noise.
[0067] In this embodiment, the quality of the data directly determines the accuracy and reliability of the final positioning system. To establish a subway tunnel environmental voiceprint database, the first step is to conduct large-scale, planned data collection to obtain high-quality, high-coverage "audio-geographic location" paired samples.
[0068] First, plan and prepare for data collection. Specifically, determine the target subway line and operating hours (e.g., morning peak, off-peak, and evening peak to cover sound characteristics under different loads), and collect and measure various information (e.g., the straight-line distance between two stations A and B within a section, the travel time of trains going up and down within the section, etc.).
[0069] Next, audio recording is performed. The recording personnel carry high-precision recording equipment and ride in engineering vehicles or normally operating subway trains to record the friction and impact sounds of wheels and rails, wind noise in subway tunnels, and other background sounds along the entire target line in sections (including both directions). The collected audio is recorded as sample audio.
[0070] It should be noted that the audio sample collection will begin from the moment the subway train departs from any station within the section (e.g., station A) and end immediately upon arrival at the next station (e.g., station B), excluding the stopping time at station A or station B.
[0071] In addition, the acquisition process should focus on collecting sounds useful to this proposal, mainly including two categories: (1) wheel-rail noise, including wheel-rail friction and impact noise, which is generated by the interaction between a specific type of wheel and a rail with a unique wear condition on a track with a fixed geometry, and has strong position specificity; (2) wind noise, i.e. tunnel aerodynamic noise, which is generated when a train travels at high speed in a tunnel, compressing the air in front and forming specific vortices. The structural characteristics of the tunnel, such as its inner diameter, wall roughness, and track bed type, directly affect the spectrum and resonance mode of this wind noise, making it a very reliable acoustic signature.
[0072] Furthermore, after collecting the sample audio, spatiotemporal alignment is required. Each sample audio segment to be collected is bound to its geographical location information, which includes at least the current geographical location, and may also include the current subway line and operating section. In addition to binding it to its geographical location, it can also be bound to the direction of travel, total travel time, and recording timestamp. For example, it might record "Audio segment A, recorded on Line X from station 1 to station 2, total travel time 75.19 seconds, recorded at 10:00 AM."
[0073] Furthermore, the collected sample audio can be preliminarily verified and backed up to ensure that the audio is clear, the location data is complete and accurate, and there is no data loss due to equipment failure. Subsequently, these sample audios with complete metadata will be securely transmitted to the server to prepare for the next step of processing.
[0074] While collecting audio samples, interference noise samples can be obtained on-site. These interference noise samples include, but are not limited to, human voices, subway announcements, door opening and closing sounds, and construction noise. Subway announcements can include platform announcements and train car announcements.
[0075] Step S20: Based on the interference noise sample, process the sample audio and extract the sample acoustic feature vector from the processed sample audio.
[0076] Then, based on the interference noise samples, the sample audio is processed, and then the acoustic feature vector is extracted from the processed sample audio, which is denoted as the sample acoustic feature vector.
[0077] In one embodiment, the sample audio is denoised based on the interference noise samples; the denoised sample audio is then segmented into frames to obtain a sample audio frame sequence; a time-frequency acoustic feature sequence is extracted based on the sample audio frame sequence; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum features; and a sample acoustic feature vector is generated by aggregating the time-frequency acoustic feature sequence.
[0078] In another embodiment, the sample audio is first subjected to noise reduction and normalization processing based on the interference noise sample to obtain standard audio; then, the standard audio is subjected to frame processing to obtain a standard audio frame sequence; based on the standard audio frame sequence, a time-frequency acoustic feature sequence is extracted; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum feature; based on the time-frequency acoustic feature sequence, the sample acoustic feature vector is aggregated to generate a sample acoustic feature vector.
[0079] The specific execution process can be found in the following examples, which will not be elaborated here.
[0080] The second implementation method performs noise reduction and normalization processing before feature extraction. Compared with the first implementation method, it can effectively suppress irrelevant noise (such as human voices and broadcasts) while preserving or even enhancing the first subway feature noise (wheel-rail friction and impact sound, wind noise) used for positioning to the maximum extent. This helps to extract more standard, pure and location-discriminative sample acoustic feature vectors.
[0081] Step S30: Divide the subway line into multiple geographic grids, and divide the sample acoustic feature vector into the corresponding geographic grids according to the geographic location information.
[0082] The system would not store an acoustic model for every infinitesimally small point, as that would be neither practical nor necessary. In this embodiment, the subway line is virtually divided into a series of continuous grids based on the track geography information. For example, each 10-meter or 20-meter interval can be defined as a grid. This method allows for the establishment of management units for subsequent acoustic models, transforming the problem of continuous geographical location into the problem of identifying a finite number of grids.
[0083] Then, based on the geographical location information corresponding to the sample acoustic feature vectors, each sample acoustic feature vector is divided into corresponding geographical grids. In this way, the originally chaotic massive sample acoustic feature vectors are organized in an orderly manner, with each grid containing several sample acoustic feature vectors recorded near that location.
[0084] Step S40: Train the sample acoustic feature vectors after segmentation to obtain the subway tunnel environment acoustic fingerprint database.
[0085] The acoustic feature vectors of the partitioned samples are used for training to obtain the acoustic fingerprint database of the subway tunnel environment. The acoustic fingerprint database of the subway tunnel environment can be constructed in different hierarchical structures.
[0086] In one embodiment, the acoustic signature database for the subway tunnel environment is constructed as a structure that includes only the grid level. Specifically, based on the acoustic feature vectors of samples divided into the same geographic grid, a grid-level acoustic model corresponding to the geographic grid is generated to form the acoustic signature database for the subway tunnel environment.
[0087] In another embodiment, the acoustic signature database for the subway tunnel environment is constructed as a two-level structure including "interval-grid". Specifically, the acoustic feature vectors of samples belonging to multiple geographic grids within the same operating interval are aggregated and trained to generate an interval-level acoustic model corresponding to the operating interval; the acoustic feature vectors of samples divided into the same geographic grid are trained to generate a grid-level acoustic model corresponding to the geographic grid; wherein, the acoustic signature database for the subway tunnel environment includes an interval-level acoustic model and a grid-level acoustic model.
[0088] In another embodiment, the acoustic signature database for subway tunnel environments is constructed as a three-tiered tower structure comprising "line-section-grid". Specifically, acoustic feature vectors of samples belonging to multiple geographic grids within the same subway line are aggregated and trained to generate a line-level acoustic model corresponding to the subway line; acoustic feature vectors of samples belonging to multiple geographic grids within the same operating section are aggregated and trained to generate a section-level acoustic model corresponding to the operating section; and a grid-level acoustic model corresponding to the geographic grid is generated based on acoustic feature vectors of samples assigned to the same geographic grid. The subway tunnel environmental acoustic signature database includes line-level acoustic models, section-level acoustic models, and grid-level acoustic models.
[0089] The subway positioning method provided in this application transforms massive amounts of chaotic sample audio into a well-structured and efficiently queryable subway tunnel environmental voiceprint database through gridding processing and statistical model training. This not only reduces the storage and computational overhead during subsequent online matching but also effectively characterizes the overall probability distribution of subway background noise through statistical models, laying a data foundation for accurate matching.
[0090] Based on any of the above embodiments Figure 3 This is the third flowchart illustrating the subway positioning method provided in this application embodiment, as shown below. Figure 3 As shown, step S20 includes: step S21, step S22, step S23 and step S24.
[0091] Step S21: Based on the interference noise samples, the sample audio is sequentially subjected to noise reduction and normalization processing to obtain standard audio.
[0092] Based on the interference noise samples, irrelevant human voices, subway broadcasts, door opening and closing sounds and / or construction noise in the sample audio are filtered out through physical filtering and / or AI (Artificial Intelligence) separation technology, while the first subway characteristic noise (wheel and rail noise and wind noise) is preserved or even enhanced, and energy normalization is performed to eliminate volume fluctuations caused by microphone gain or distance, and finally the standard audio is obtained.
[0093] Step S22: Perform frame segmentation on the standard audio to obtain a standard audio frame sequence.
[0094] Framing is a fundamental operation that divides a continuous audio signal stream into a series of short time intervals for analysis. It is based on the assumption that the spectral characteristics of a sound signal are basically stable within a short period of time (usually 20-40 milliseconds), i.e., the "short-time stationarity" hypothesis.
[0095] During the framing process, when selecting the frame length, it is necessary to consider that the frame length is sufficient to accommodate several periodic pulses of wheel-rail impacts and capture the stable spectrum of wind noise, while also meeting the requirements of short-term stability. In one embodiment, the frame length is set to 20-30 milliseconds, preferably 25 milliseconds.
[0096] Frame shift determines the duration of the overlap between adjacent frames. This overlap strategy ensures signal continuity and avoids the loss of any transient, important acoustic events (such as brief track gap impacts) due to framing, thus providing continuous and sufficient temporal context for subsequent feature extraction. In one implementation, the frame shift is set to 10 milliseconds. When the frame length is 25 milliseconds, there is a 15-millisecond overlap between adjacent frames.
[0097] Step S23: Extract features from the standard audio frame sequence to obtain a time-frequency acoustic feature sequence; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum feature.
[0098] Then, feature extraction is performed on each audio frame in the standard audio frame sequence to obtain the time-frequency acoustic features of each audio frame, thus forming a time-frequency acoustic feature sequence.
[0099] When extracting time-frequency acoustic features, at least one of Mel-frequency cepstral coefficients (MFCCs) feature vectors and Mel-spectral features can be extracted. In one embodiment, MFCCs can be used as the core feature, supplemented by Mel-spectral graphs to support more complex deep learning models, thereby constructing a robust and highly discriminative acoustic signature database for subway tunnel environments.
[0100] Furthermore, the Mel frequency cepstral coefficient eigenvectors are obtained by sequentially performing pre-emphasis, short-time Fourier transform and power spectrum calculation, Mel filter bank mapping, and logarithmic compression and discrete cosine transform on each audio frame in the standard audio frame sequence.
[0101] The specific steps for extracting the eigenvectors of Mel frequency cepstral coefficients are as follows: The first step is the pre-emphasis process, which involves applying a first-order high-pass filter (typically with a transfer function of H(z) = 1 - 0.97z) to each audio frame signal. -1 This step aims to enhance high-frequency components. It compensates for the high-frequency attenuation of the signal during propagation, making the high-frequency details generated by wheel-rail friction and impact more prominent.
[0102] The next step is the short-time Fourier transform and power spectrum calculation. This step involves performing a short-time Fourier transform on the pre-emphasized signal of each audio frame, converting it from the time domain to the frequency domain to obtain a linear spectrum, and then calculating its linear power spectrum.
[0103] Next, Mel filter bank mapping is performed, whereby the linear power spectrum is passed through a Mel-scale triangular filter bank to distort the linear spectrum to a Mel frequency domain that better matches human auditory perception. By integrating the energy within the filter bank, the spectral characteristics are smoothed and dimensionality reduced, effectively weakening harmonic details while preserving the crucial spectral envelope shape. It should be noted that if the number of filters is too small, the bandwidth of each filter will be too wide, leading to a significant loss of spectral details. For subway noise, the fine harmonic structure of wheel-rail friction sound and the resonance peaks of tunnel wind noise may be indistinguishable because they fall within the same wide bandwidth, making the acoustic signatures of different sections similar and reducing positioning accuracy. If the number of filters is too large, the bandwidth of each filter will be very narrow. While this can capture finer spectral details, it will cause problems such as excessive dimensionality and decreased stability, ultimately reducing the robustness of the system. Therefore, in this invention, a filter bank containing 20 to 40 filters with a center frequency evenly distributed according to the Mel frequency scale is used, such as the traditional 26-group filter bank, i.e., 12-dimensional MFCC + 1-dimensional energy.
[0104] Finally, logarithmic compression and discrete cosine transform (DCT) are performed. This step takes the logarithm of the output energy of each filter bank. This operation not only conforms to the logarithmic response of the human ear to sound intensity but also reduces the correlation between feature dimensions. The resulting logarithmic energy sequence is then subjected to a discrete cosine transform, retaining the first 13 coefficients (derived from the aforementioned filter banks, i.e., the 12-dimensional MFCC values and the 0th coefficient, the energy value). These coefficients constitute the final MFCC feature vector, which centrally represents the core features of the audio spectral envelope of that frame.
[0105] To further enhance the robustness of voiceprint features, the extracted MFCC feature vectors can be post-processed. For example, the first-order difference (Delta) coefficients and second-order difference (Delta-Delta) coefficients can be calculated to capture the dynamic changes in the spectrum, ultimately forming a set of composite feature vectors with higher dimensions and stronger representation capabilities.
[0106] Further, the specific steps for extracting Mel spectral features are as follows: the linear spectrum obtained from the short-time Fourier transform is directly passed through the aforementioned Mel filter bank, and the output energy of each filter is converted into logarithmic energy. The value of the logarithmic energy is then filled into the position of the corresponding frame and filter (i.e., Mel band) in the two-dimensional matrix. The final sample audio is represented as a two-dimensional matrix (time frame × Mel band), i.e., a Mel spectrogram. It can be viewed as an acoustic image, where the color intensity represents the energy strength, the horizontal axis is time, and the vertical axis is Mel frequency. When used as input to a machine learning model, it becomes the Mel spectral feature. This feature preserves the dynamic information of sound changes over time, making it particularly suitable for inputting into convolutional neural networks for end-to-end feature learning and location recognition.
[0107] Step S24: Based on the time-frequency acoustic feature sequence, aggregate to generate the sample acoustic feature vector.
[0108] To facilitate storage and matching, the time-frequency acoustic features of all audio frames within a sample audio are typically statistically aggregated, for example, by calculating their mean and standard deviation in each dimension, thus forming a more representative statistical feature vector, denoted as the sample acoustic feature vector.
[0109] Furthermore, the aggregated sample acoustic feature vectors are associated with the geographical location information corresponding to the sample audio.
[0110] The subway positioning method provided in this application, through the above feature extraction process, transforms sample audio from the subway tunnel environment into low-dimensional, compact digital features that can be efficiently processed by machine learning models, providing standardized input for geographic gridding and model training in the subsequent construction of the subway tunnel environment voiceprint database.
[0111] Based on any of the above embodiments, step S21 includes: step S211, step S212, step S213, step S214, step S215 and step S216.
[0112] It should be noted that the execution order of steps S211 and S212-S213 is not important and they can be executed in parallel.
[0113] Step S211: Perform time-frequency conversion on the sample audio to obtain the sample time-frequency spectrum.
[0114] The sample audio is converted from time to frequency to obtain a time-spectrum, which is denoted as the sample time-spectrum.
[0115] Step S212: Identify the first random interference noise in the sample audio based on the interference noise sample, and identify the first subway characteristic noise in the sample audio.
[0116] Step S213: Filter the first random interference noise and enhance the target frequency band of the first subway characteristic noise to extract the target subway characteristic parameters.
[0117] Based on the interference noise samples, random interference noise (denoted as the first random interference noise) in the sample audio is identified, and subway characteristic noise (denoted as the first subway characteristic noise) in the sample audio is also identified. The first subway characteristic noise includes at least one of wheel-rail noise and wind noise, and wheel-rail noise includes wheel-rail friction and impact sounds; the first random interference noise includes, but is not limited to, human voices, subway announcements, door opening and closing sounds, and construction noise.
[0118] At the physical level, for the first type of random interference noise, taking human voice and subway broadcasts as examples, since human voice and subway broadcasts have a fundamental frequency (85-255Hz) and harmonic structure, we first calculate the higher-order statistics of the audio frame (such as the third-order cumulant). Utilizing the difference between the non-Gaussian nature of human voice / subway broadcasts and the near-Gaussian nature of wheel-rail / wind noise, we design a zero-phase IIR (Infinite Impulse Response) notch filter to accurately remove the harmonic frequency band while preserving the core characteristics of adjacent frequency bands.
[0119] At the physical level, based on the physical characteristics of wheel-rail noise and wind noise, joint filtering is performed on the "frequency band-characteristics".
[0120] To address the mid-to-high frequency (2-8kHz) pulse characteristics of wheel-rail noise, a harmonic comb filter is applied to enhance periodic impacts using a kurtosis index. Specifically, a pulse enhancement filter is employed: first, pulse segments in the sample audio are detected using a kurtosis index (e.g., kurtosis > 5); then, a harmonic comb filter is used, with its tooth pitch dynamically adjusted based on the known track gap spacing and train speed, thereby specifically enhancing the periodic impact components generated by the track gaps.
[0121] To address wind noise, an adaptive line enhancer is applied in the 200-800Hz frequency band to lock the resonant main peak determined by the Strouhal number. Specifically, using an adaptive line enhancer with the 200-800Hz frequency band as the desired signal, the dominant frequency component of the wind noise is iteratively tracked and enhanced using the LMS (Least Mean Square) algorithm.
[0122] These enhanced physical parameters, such as the impulse index and the Strauhall peak value, constitute the characteristic parameters of the target metro.
[0123] Furthermore, dual verification and fail-safe mechanisms can be employed to ensure reliability.
[0124] The first step is to perform a physical consistency check, which involves recalculating the spectral centroid and kurtosis of the physically filtered audio signal. If the centroid drift exceeds ±10% or the kurtosis decreases by more than 30%, it is considered over-filtered, and the system will automatically revert to a conservative lightweight spectral subtraction. Specifically, in audio segments identified as pure interference by a CNN (Convolutional Neural Network), the noise power spectrum P is estimated. noise(f) Subsequently, from the noisy speech spectrum P corresponding to the sample audio... signal(f) Subtract from the middle, i.e., P enhanced(f) =P signal(f) -α*P noise(f) , where P enhanced(f) This represents the power spectrum after denoising, where α is an over-subtraction factor that is dynamically adjusted based on the confidence level of the CNN output.
[0125] Secondly, a signal-to-noise ratio (SNR) threshold is set. If the estimated global SNR is lower than 0 dB (indicating that the environment is too noisy), the feature extraction will be abandoned and the noisy environment will be reported, requesting a delay in sampling to avoid low-quality features from being included in the database or participating in matching.
[0126] Step S214: Input the target subway feature parameters and the sample time spectrum into a preset neural network model to obtain the feature mask output by the preset neural network model.
[0127] To achieve further fine separation, a lightweight neural network model can be deployed to enable fine separation of the "mask-reconstruction" process.
[0128] The pre-defined neural network model is trained as follows: First, an interference sample set is established through field data collection. Specifically, ambient sounds without wheel-rail movement are recorded in stationary scenarios such as trains, platforms, tunnels, and maintenance depots to systematically capture steady-state or transient interference noise such as air conditioning, fan noise, electrical noise, and typical broadcasts. These pure interference samples are manually labeled and clustered, for example, by extracting features such as MFCC, spectral centroid, and zero-crossing rate. Based on the manually labeled results and the clustered sample feature vectors, an interference sample set is constructed. Then, a lightweight CNN (Convolutional Neural Network) classification network, such as 1-D MobileNet, is trained based on this interference sample set, resulting in the trained neural network model, i.e., the pre-defined neural network model. This pre-defined neural network model can quickly determine whether there is interference noise that needs to be filtered out based on a 0.5-second audio clip and provides prior knowledge for online processing.
[0129] Furthermore, during the training process of the preset neural network model, the loss function used is specially designed, which simultaneously constrains the reconstruction error of the feature preservation region and the energy of the interference suppression region, and introduces kurtosis loss to ensure that the pulse characteristics of the wheel and rail are not excessively smoothed.
[0130] The sample time-frequency spectrum is input into a pre-defined neural network model, which extracts deep time-frequency features from the spectrum to guide mask generation. Simultaneously, a core feature prior branch is introduced into the bottleneck layer of the pre-defined neural network model. The target subway feature parameters (including the wheel-rail impulse index and the Strauhall peak value of wind noise) are input as additional guiding information, enabling the model to focus more on regions with significant physical characteristics. Finally, the pre-defined neural network model outputs a binary mask, denoted as the feature mask, which identifies the probability that each time-frequency unit belongs to the core feature.
[0131] Step S215: Based on the feature mask, perform weighted reconstruction and time-domain transformation on the sample time-spectrum to obtain the processed sample audio.
[0132] Finally, the sample time-frequency spectrum is reconstructed using this feature mask to obtain a clean core feature waveform, which is then transformed in the time domain to obtain the processed sample audio.
[0133] Step S216: Perform global level normalization and spectral tilt compensation on the processed sample audio in sequence to obtain the standard audio.
[0134] The purpose of normalization is to eliminate the overall volume fluctuation of the audio signal caused by factors such as differences in recording equipment gain and different distances between the microphone and the sound source, so as to ensure that the extracted voiceprint features only reflect the essential spectral shape of the sound, rather than irrelevant energy levels.
[0135] In this embodiment, a two-level normalization strategy (including global level normalization and spectrum tilt compensation) is adopted to achieve fine energy normalization.
[0136] During global level normalization, the root mean square energy of the processed sample audio is calculated over the entire time domain. Then, the energy of all samples is scaled to a uniform target level, thereby eliminating the huge volume differences between different recording segments and providing a consistent starting point for subsequent processing.
[0137] In the spectrum tilt compensation process, the background noise of subways, especially wheel-rail friction noise, may have a specific tilt in its spectrum. Simple global scaling may not be able to completely eliminate this spectral tilt. Therefore, this application introduces an online spectrum shifting technique to calculate the spectrum of the processed sample audio and perform a first-order linear fit on it in the Mel frequency domain. By subtracting this fitted tilt line, the spectrum tilt caused by some equipment frequency response or specific acoustic environment can be compensated, making the finally extracted MFCC and other features more focused on key details such as local peaks and valleys in the spectrum, further improving the feature's ability to distinguish different locations.
[0138] The subway positioning method provided in this application uses the aforementioned physical filtering and AI separation techniques to filter out irrelevant human voices / broadcasts from sample audio, while retaining or even enhancing the first subway characteristic noise. This achieves targeted feature separation, ensuring the highest purity of the subsequently extracted voiceprint features. Furthermore, through the aforementioned normalization process, it ensures that audio collected from different devices, at different times, and in different carriage locations, as long as their acoustic nature is the same, can be mapped to highly similar feature vector spaces.
[0139] Based on any of the above embodiments, step S40 includes: step S41, step S42 and step S43.
[0140] Step S41: Aggregate the sample acoustic feature vectors of multiple geographic grids belonging to the same subway line for training, and generate the line-level acoustic model corresponding to the subway line.
[0141] To achieve efficient and accurate voiceprint matching, this embodiment of the invention adopts a hierarchical acoustic model architecture, constructing the subway tunnel environment voiceprint database into a three-level tower structure of "line-section-grid".
[0142] For each independent metro line (such as "Line 1" and "Line 2"), a massive number of sample acoustic feature vectors are sampled from all geographic grids belonging to that metro line to form a hybrid dataset representing the global acoustic characteristics of that metro line. A line-level acoustic model is trained based on this dataset. Specifically, by performing statistical analysis on the sample acoustic feature vectors of the entire line, such as principal component analysis (PCA) or cluster analysis, features that can represent the global acoustic background of the line are extracted. These features may include, but are not limited to: (1) overall spectral tone: due to differences in vehicle models, tunnel construction standards, and average operating speeds, different lines will exhibit unique overall spectral energy distribution profiles; (2) characteristic frequency band distribution: some lines may have sustained high energy in specific frequency bands (such as 500-800Hz) due to their track structure or environmental reasons; (3) statistical regularity of acoustic events: the statistical regularity of periodic impact sounds caused by factors such as track gap density and number of curves varies on different lines. Based on the above aggregated features, a line-level acoustic model is trained for each metro line. This model does not capture the sound of a specific location, but rather the unique auditory character of the entire subway line caused by its inherent properties (such as the type of train used, tunnel construction standards, and average operating speed). For example, the overall spectral energy distribution and reverberation characteristics of the background noise of a fully underground line will show systematic differences from those of a line that includes surface sections.
[0143] In one implementation, the line-level acoustic model can be a Gaussian mixture model used to characterize the overall probability distribution of the acoustic features of the subway line. This model can describe the probability of hearing a specific sound (such as wheel-rail noise of a specific frequency) at the location of the grid, and what the typical characteristics of this sound are. It captures the overall statistical distribution characteristics of the background noise at that location.
[0144] In another implementation, the line-level acoustic model can be a deep autoencoder whose bottleneck layer output can serve as the acoustic summary vector for the line. During training, a deep metric learning approach is used to train an embedding network that maps multiple feature vectors of each grid to points close to each other in a shared embedding space, thereby learning a more discriminative grid-based deep acoustic signature code.
[0145] Step S42: Aggregate the sample acoustic feature vectors of multiple geographic grids belonging to the same driving section for training, and generate the interval-level acoustic model corresponding to the driving section.
[0146] Each train section (such as a tunnel section between two adjacent stations) has a unique acoustic background due to its relatively consistent physical structure, including tunnel radius, curves, and gradients. Therefore, the acoustic feature vectors of samples from all geographic grids belonging to the same train section are aggregated, and a section-level acoustic model is trained for this dataset. This section-level acoustic model captures the relatively uniform acoustic features that distinguish the entire train section from other train sections on the same line, providing a crucial basis for fine-grained section-level matching in hierarchical matching.
[0147] Furthermore, the interval-level acoustic model can use a Gaussian Mixture Model (GMM) to characterize the overall probability distribution of the sample acoustic feature vectors in that interval.
[0148] Step S43: Train the model based on the acoustic feature vectors of samples divided into the same geographic grid to generate a grid-level acoustic model corresponding to the geographic grid.
[0149] The subway tunnel environmental acoustic model database includes the line-level acoustic model, the section-level acoustic model, and the grid-level acoustic model.
[0150] At the finest granularity, a grid-level acoustic model is trained for each geographic grid. This grid-level acoustic model is trained based on the sound feature vectors of all samples divided into that grid, and is designed to accurately describe the statistical distribution characteristics of background noise in that specific geographic grid.
[0151] The subway positioning method provided in this application constructs a subway tunnel environmental voiceprint database with a three-level tower model structure of "line-section-grid" as described above. This hierarchical design not only truly reflects the physical topology of the subway network, but also transforms the flat query of massive grid points into a hierarchical and efficient retrieval, providing a structured data foundation for subsequent hierarchical matching algorithms.
[0152] Based on any of the above embodiments, the subway tunnel environmental acoustic model database includes a line-level acoustic model, a section-level acoustic model, and a grid-level acoustic model. Figure 4 This is the fourth flowchart illustrating the subway positioning method provided in the embodiments of this application, as follows: Figure 4 As shown, step S120 includes: step S121, step S122 and step S123.
[0153] Step S121: Match the target acoustic feature vector with the line-level acoustic model to determine the candidate line set.
[0154] First, the target acoustic feature vector is matched with the line-level acoustic model to determine the candidate line set.
[0155] Specifically, the target acoustic feature vector is matched with the line-level acoustic models (such as line-level GMMs or acoustic summary vectors) of each subway line in the subway tunnel environment acoustic fingerprint database. Based on the first matching result, the subway lines that meet the first preset conditions are selected as the candidate line set.
[0156] In one implementation, if the line-level acoustic model is constructed using a Gaussian mixture model, log-likelihood is preferred as the similarity matching criterion. Specifically, for each line-level acoustic model, the system calculates the expected log-likelihood of the probability distribution of the target acoustic feature vector belonging to that model, and takes the average value as the final voiceprint matching score. This voiceprint matching score quantitatively reflects the statistical degree of similarity between the ambient sound currently collected by the terminal and the sound template at a specific location in the library.
[0157] In another implementation, if the line-level acoustic model is constructed using other models (such as a deep autoencoder), cosine similarity or Mahalanobis distance can also be used to characterize the voiceprint matching score.
[0158] Furthermore, the setting of the first preset condition includes, but is not limited to: (1) the voiceprint matching score is greater than the first preset score threshold; (2) the top N lines with the largest voiceprint matching scores; (3) dynamic strategy: prioritize the selection of lines with voiceprint matching scores exceeding the first preset score threshold as high confidence candidates; if the number of selected lines is too large, further select the top-ranked specific number (such as the top N) entries; if the number of selected lines is insufficient, relax the restrictions to ensure that a sufficient number of candidate lines flow into the next processing stage.
[0159] Using the above methods, the search scope can be quickly narrowed down from the entire subway network to a few lines.
[0160] Step S122: Match the target acoustic feature vector with the interval-level acoustic model corresponding to the candidate line set to determine the candidate interval set.
[0161] Then, within the selected candidate route set, the target acoustic feature vector is further matched with the interval-level acoustic model of each driving section under the corresponding route of the candidate route set. Based on the second matching result, on each candidate route, the driving section that meets the second preset condition is selected as the candidate interval set.
[0162] Similarly, when performing similarity matching, the voiceprint matching score can be characterized by calculating the expected value of log-likelihood, cosine similarity, or Mahalanobis distance.
[0163] The second preset condition can be set in the following ways: (1) the voiceprint matching score is greater than the second preset score threshold; (2) the top M intervals with the highest voiceprint matching scores; (3) dynamic strategy: prioritize the selection of intervals with voiceprint matching scores exceeding the second preset score threshold as high confidence candidates; if too many are selected, further select the top-ranked specific number (such as the top M) entries; if the number of selected is insufficient, relax the restrictions to ensure that a sufficient number of candidate intervals flow into the next processing stage.
[0164] By using the above method, the search scope can be narrowed down from the entire subway line to a few key continuous tunnel sections.
[0165] Step S123: Match the acoustic feature vector with the grid-level acoustic model corresponding to the candidate interval set to determine the candidate location set.
[0166] Next, within the selected candidate interval set, the final similarity matching is performed on the finest-grained grid-level acoustic model. Based on the third matching result, grids that meet the third preset conditions are selected as candidate location sets for each candidate interval.
[0167] Similarly, when performing similarity matching, the voiceprint matching score can be characterized by calculating the expected value of log-likelihood, cosine similarity, or Mahalanobis distance.
[0168] The third preset condition can be set in the following ways, including but not limited to: (1) the voiceprint matching score is greater than the third preset score threshold; (2) the top L grids with the highest voiceprint matching scores; (3) dynamic strategy: prioritize the selection of grids with voiceprint matching scores exceeding the third preset score threshold as high confidence candidates; if too many grids are selected, further select the top-ranked specific number (such as the top L) entries; if the number of grids selected is insufficient, relax the restrictions to ensure that a sufficient number of candidate grids flow into the next processing stage.
[0169] After three layers of screening, a controllable set of candidate locations is finally generated. Each candidate location in this set includes: a geographic grid ID, the corresponding latitude and longitude coordinates, and a voiceprint matching score calculated at this level.
[0170] The subway positioning method provided in this application adopts a three-level hierarchical matching strategy of "line → section → grid", which aims to filter layer by layer while ensuring accuracy, thereby greatly improving search efficiency.
[0171] Based on any of the above embodiments Figure 5 This is the fifth flowchart illustrating the subway positioning method provided in the embodiments of this application, as shown below. Figure 5As shown, step S130 includes: step S131, step S132, step S133 and step S134.
[0172] Step S131: Obtain the historical location and first timestamp of the target object.
[0173] Step S132: Based on the second timestamp corresponding to the target acoustic feature vector, the historical positioning location, and the preset maximum running speed of the first timestamp, determine the theoretically achievable range.
[0174] To ensure the reliability of the results, the target location is further determined by combining temporal context filtering with the candidate location set.
[0175] First, obtain the target object's historical location and first timestamp. The historical location is the location of the last successful location, and the first timestamp is the time corresponding to the last successful location.
[0176] Then, based on the second and first timestamps corresponding to the target acoustic feature vector and the preset maximum running speed, the theoretical maximum moving distance is calculated. The specific calculation formula is as follows: D max =V max *(T current -T prev ); Among them, D max V represents the theoretical maximum distance traveled. max T represents the preset maximum operating speed of the subway train. current Indicates the second timestamp, T prev Indicates the first timestamp.
[0177] Based on historical location and theoretical maximum movement distance, the theoretically reachable range is determined. Specifically, a range is defined on the map starting from the historical location P. prev为 Center, theoretical maximum movement distance D max The region is defined by radius , and the theoretically reachable range is defined by radius .
[0178] Step S133: Based on the theoretically reachable range, the candidate position set is filtered to obtain the remaining candidate positions.
[0179] Step S134: Determine the target location based on the remaining candidate locations.
[0180] By traversing through each candidate position in the candidate position set, candidate positions that are outside the theoretically reachable range are directly eliminated to obtain the remaining candidate positions. This step can effectively eliminate "ghost" positions caused by brief similarity of sound.
[0181] Then, based on the remaining candidate locations, the target location is determined.
[0182] In one embodiment, the candidate position with the highest voiceprint matching score among the remaining candidate positions is determined as the target location of the target object.
[0183] In another embodiment, operational constraint information of the subway environment where the target object is located is obtained; the operational constraint information includes at least one of path topology, speed limit rules, and direction of travel; based on the operational constraint information, the voiceprint matching scores corresponding to the remaining candidate positions are weighted to obtain a comprehensive confidence score; based on the comprehensive confidence score, the target location is determined from the remaining candidate positions. The specific execution process can be referred to in the following embodiments, and will not be elaborated here.
[0184] The subway positioning method provided in this application, based on a candidate location set, further combines temporal context filtering to determine the target location. Specifically, it utilizes the physical motion limit principle of the real world to determine the theoretically reachable range, thereby filtering the candidate location set. This effectively eliminates "ghost" locations that are extremely far away due to brief similarity in sound, significantly improving the spatiotemporal rationality of the subway positioning results.
[0185] Based on any of the above embodiments, step S134 includes: step S1341, step S1342 and step S1343.
[0186] Step S1341: Obtain the operation constraint information of the subway environment where the target object is located; the operation constraint information includes at least one of path topology, speed limit rules and direction of travel.
[0187] In this embodiment, based on the temporal context filtering, further execution logic verification and weighting are performed to transform the operation rules of the subway network into quantified weighting factors for the final confidence assessment.
[0188] The built-in digital model of the subway line is invoked to obtain the operational constraint information of the subway environment in which the target object is located; the operational constraint information includes at least one of the following: path topology, speed limit rules, and direction of travel.
[0189] Step S1342: Based on the running constraint information, the voiceprint matching scores corresponding to the remaining candidate positions are weighted to obtain a comprehensive confidence score.
[0190] A logical rationality weighting factor is determined based on operational constraint information. The calculation of this logical rationality weighting factor is based on at least one of the following evaluation rules: (1) Directional consistency weighting: If candidate position C i The track located directly in front of the train's direction of travel is assigned a higher directional weight W.direction For example, W direction =1.2; if located on the opposite track, a significantly lower directional weight is assigned, for example, W direction =0.3).
[0191] (2) Location type weighting: Considering that the acoustic environment inside the tunnel is usually more stable than that of the platform, if candidate location C i If it is located within a tunnel section, then its location type weight W type Slightly higher, for example, W type =1.1; If located in the platform area, the location type weight is slightly lower, for example, W type =0.9.
[0192] (3) Speed consistency verification: based on historical position P prev With candidate position C i The average velocity V is calculated based on the distance and the time difference between the first and second timestamps. avg If V avg The speed limit V for this section was not exceeded. limit Then the speed weight W speed =1.0; if V avg Significantly exceeds V limit Then W speed Exponential decay occurs when the ratio exceeds the limit, for example, W speed =exp(-(V avg -V limit ) / V limit ).
[0193] Candidate position C i The logical rationality weighting factor W logic-i It is the product of its various sub-weights. If the above three sub-weights are included, then W logic-i =W direction *W type *W speed .
[0194] Then, the voiceprint matching scores corresponding to the remaining candidate positions are weighted using this logical rationality weighting factor to obtain the comprehensive confidence score. That is, for each remaining candidate position, the comprehensive confidence score = voiceprint matching score * logical rationality weighting factor.
[0195] Step S1343: Based on the comprehensive confidence score, determine the target location from the remaining candidate locations.
[0196] Finally, from the remaining candidate positions, the one with the highest overall confidence score is selected as the target location.
[0197] Furthermore, to ensure the reliability of the positioning results, the system performs result verification. The verification methods include, but are not limited to: (1) Verification based on an absolute confidence threshold. That is, an absolute confidence threshold is set. If the overall confidence score is greater than the absolute confidence threshold, the positioning result is determined to be of high confidence; otherwise, it is of low confidence. (2) Verification based on a relative win ratio. The ratio of the highest score to the second highest score in the overall confidence score is calculated and recorded as the relative win ratio. If the highest score is significantly higher than the second highest score, that is, the relative win ratio is greater than the preset win ratio threshold (e.g., 3), the positioning result is determined to be of high confidence; otherwise, it is of low confidence. (3) Verification based on both the absolute confidence threshold and the relative win ratio. If the overall confidence score is greater than the absolute confidence threshold and the relative win ratio is greater than the preset win ratio threshold, the positioning result is determined to be of high confidence; otherwise, it is of low confidence.
[0198] It should be noted that setting a high win ratio threshold is to ensure that the optimal candidate position is significantly better than other candidate positions in terms of acoustic features and logical rationality, thereby greatly reducing the risk of false matching due to similar local features.
[0199] Furthermore, if the target is determined to have a high confidence level, the target location will be sent to the mobile terminal.
[0200] Furthermore, if the result is determined to be of low confidence, a conservative strategy will be adopted, such as discarding the current result, not updating the terminal location, or immediately triggering the mobile terminal to conduct a new round of sampling and reporting.
[0201] The subway positioning method provided in this application breaks through the limitations of simply relying on acoustic similarity. It transforms the strong operational rules of subways, such as one-way travel and speed limits, into a quantified digital penalty mechanism. This multimodal information fusion decision greatly eliminates matching ambiguities caused by the similarity of sounds from adjacent tracks, and provides extremely high robustness and accuracy, especially in solving pain points such as riding in the wrong direction.
[0202] Figure 6 This is the sixth flowchart illustrating the subway positioning method provided in this application embodiment, as shown below. Figure 6 As shown, the subway positioning method includes steps S210, S220 and S230.
[0203] Step S210: Obtain the target acoustic feature vector of the subway environment audio where the target object is located.
[0204] The subway positioning method provided in this application is applied to a mobile terminal.
[0205] When a target carrying a mobile terminal enters a subway station or subway car, the terminal uses its built-in microphone to collect audio from the subway environment in that location.
[0206] After acquiring the audio, the mobile terminal does not directly upload the lengthy raw audio; instead, it performs real-time analysis locally. Specifically, a feature extraction algorithm runs locally to convert the acquired analog signal, which includes characteristic subway noises such as wheel-rail noise and tunnel wind noise, into a digitized acoustic feature vector, denoted as the target acoustic feature vector. This process significantly reduces the amount of data transmitted over the network and fundamentally protects passengers' privacy during conversations.
[0207] Step S220: Send the target acoustic feature vector to the server.
[0208] Then, the mobile terminal extracts the target acoustic feature vector and sends it to the remote server via the mobile network.
[0209] Alternatively, the current timestamp, device ID, and other information can be packaged together with the target acoustic feature vector and sent to the server.
[0210] Step S230: Receive the target location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic signature database; wherein, the subway tunnel environment acoustic signature database is constructed based on first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
[0211] After obtaining the voiceprint feature vector, the server performs a rapid search in the pre-established subway tunnel environment voiceprint database to match the target acoustic feature vector with the subway tunnel environment voiceprint database, thereby determining the candidate location set corresponding to the target object, and then determining the target location of the target object based on the candidate location set.
[0212] Correspondingly, the mobile terminal receives the target location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database.
[0213] The metro tunnel environmental acoustic signature database is a geolocation-acoustic model mapping database constructed based on the first characteristic noise of the metro collected within the metro tunnel. The first characteristic noise of the metro includes at least one of wheel-rail noise and wind noise. Wheel-rail noise refers to the friction / impact sound generated by the specific wear state of the wheels and rails, while wind noise refers to the aerodynamic wind noise formed by the influence of tunnel diameter and wall roughness.
[0214] The subway positioning method provided in this application involves acquiring the target acoustic feature vector of the subway environment audio at the location of the target object via a mobile terminal; then, sending the target acoustic feature vector to a server; and subsequently receiving the target positioning location returned by the server based on the target acoustic feature vector and a subway tunnel environment acoustic signature database. The subway tunnel environment acoustic signature database is constructed based on first subway characteristic noise collected within the subway tunnel, which includes at least one of wheel-rail noise and wind noise. Through this method, the mobile terminal only needs to acquire the acoustic feature vector of the subway environment audio at the location of the target object, without relying on high-precision inertial navigation devices or specific radio signal scanning modules. Environmental data acquisition can be completed using the mobile terminal's existing microphone, solving the technical deficiency of ordinary mobile terminals that cannot maintain accurate positioning over long distances due to low-precision inertial devices. Simultaneously, the mobile terminal acquires the acoustic feature vector locally and sends it to the server, which performs acoustic signature database matching based on the subway tunnel environment acoustic signature database. Subsequently, the mobile terminal directly receives the high-precision target positioning location. This mode avoids the problem of mobile terminals needing long periods of continuous sampling to converge due to radio frequency multipath effects in enclosed metal carriages. It utilizes acoustic characteristics to achieve rapid response under extremely short acquisition windows. At the same time, it offloads the core matching computing power to the server, ensuring that the terminal device can continuously and stably obtain location updates in the subway carriage without increasing the additional computing burden on the mobile terminal.
[0215] Based on any of the above embodiments, step S210 includes: step S211, step S212, step S213 and step S214.
[0216] Step S211: Collect the audio of the current subway environment.
[0217] Collect audio of the current subway environment.
[0218] Step S212: Identify the second random interference noise, fixed equipment noise, and second subway characteristic noise in the current subway environment audio.
[0219] Identify random interference noise (denoted as second random interference noise), fixed equipment noise, and subway characteristic noise (denoted as second subway characteristic noise) in the current subway environment audio.
[0220] The second type of subway characteristic noise includes at least one of wheel-rail noise and wind noise. Wheel-rail noise includes wheel-rail friction and impact noise. Fixed equipment noise includes friction noise at vehicle connection points, air conditioning operation noise, etc. The second type of random interference noise includes at least one of human voices, subway broadcasts, door opening and closing sounds, and construction noise.
[0221] The process for identifying stationary equipment noise is as follows: Calculate the PSD (Power Spectral Density) of the audio frame of the current subway environment audio, and perform similarity calculation (e.g., cosine similarity calculation) with a pre-stored stationary equipment noise template. When the similarity exceeds a threshold, the presence of this type of noise is determined.
[0222] The process of obtaining the fixed equipment noise template is as follows: In the offline stage, pure fixed equipment noise is recorded in scenarios such as stationary subway trains and maintenance depots, and its typical average PSD is extracted as the fixed equipment noise template.
[0223] Step S213: Filter the second random interference noise and the fixed equipment noise, and enhance the target frequency band of the second subway characteristic noise to obtain the target audio.
[0224] To address the second type of random interference noise, taking human voice and subway broadcasts as examples, which have a fundamental frequency (85-255Hz) and harmonic structure, we first calculate the higher-order statistics of the audio frames (such as the third-order cumulant). Utilizing the difference between the non-Gaussian nature of human voice / broadcast and the near-Gaussian nature of wheel / rail / wind noise, we design a zero-phase IIR notch filter to accurately remove the harmonic frequency band while preserving the core characteristics of adjacent frequency bands.
[0225] For stationary equipment noise, spectral tilt compensation and selective attenuation are applied. Specifically, the system identifies specific frequency bands where the noise energy of stationary equipment is concentrated (such as the low-frequency hum of an air conditioner) and applies a wide and shallow notch filter to these bands, rather than global attenuation, to minimize the impact on the core characteristic frequency bands.
[0226] For the characteristic noise of the second subway, a joint filtering method is used based on the physical characteristics of wheel-rail noise and wind noise, targeting the "frequency band-characteristics".
[0227] To address the mid-to-high frequency (2-8kHz) pulse characteristics of wheel-rail noise, a harmonic comb filter is applied to enhance periodic impacts using a kurtosis index. Specifically, a pulse enhancement filter is employed: first, pulse segments in the sample audio are detected using a kurtosis index (e.g., kurtosis > 5); then, a harmonic comb filter is used, with its tooth pitch dynamically adjusted based on the known track gap spacing and train speed, thereby specifically enhancing the periodic impact components generated by the track gaps.
[0228] To address wind noise, an adaptive line enhancer is applied in the 200-800Hz frequency band to lock the resonant peak determined by the Strouhal number. Specifically, using an adaptive line enhancer with the 200-800Hz frequency band as the desired signal, the dominant frequency component of the wind noise is iteratively tracked and enhanced using the LMS algorithm.
[0229] The current subway environment audio that has undergone the above filtering and enhancement processes is denoted as the target audio.
[0230] Step S214: Extract the target acoustic feature vector based on the target audio.
[0231] Based on the target audio, extract the target acoustic feature vector.
[0232] Specifically, the target audio undergoes global level normalization and spectral tilt compensation. Then, the normalized target audio is segmented into frames to obtain a target audio frame sequence. Feature extraction is performed on the target audio frame sequence to obtain a time-frequency acoustic feature sequence, denoted as the target time-frequency acoustic feature sequence. The target time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vectors and Mel spectral features. Based on the target time-frequency acoustic feature sequence, a target acoustic feature vector is generated. The specific execution process can be referred to the above sample acoustic feature vector extraction process, which will not be elaborated here.
[0233] The subway positioning method provided in this application embodiment can not only filter out random interference noise such as human voices and fixed equipment noise with extremely high efficiency, but also maximize the preservation of the details of subway characteristic noise on terminals with limited computing power. This greatly improves the accuracy of model matching on the subsequent server side and enhances the positioning robustness of the system in extremely noisy environments.
[0234] Based on any of the above embodiments, step S210 includes: step S215, step S216 and step S217.
[0235] Step S215: In the first working mode, monitor the target sound events in the subway environment where the target object is located.
[0236] Step S216: When the target sound event is detected, switch from the first working mode to the second working mode.
[0237] The first working mode, also known as the normal mode, is a low-power sleep monitoring state. It does not perform full-link recording or complex feature extraction, but only runs a lightweight sound event detection algorithm to monitor target sound events, thereby achieving macroscopic perception and rough positioning of the train's operating status, thus saving terminal power and achieving continuous monitoring.
[0238] The second working mode, also known as the precision mode, refers to a high-power, high-precision, and short-lived operating state in which the mobile terminal temporarily initiates a complex full-link acoustic signal processing and cloud interaction process in order to obtain high-precision absolute position information and complete the unique confirmation of the train's travel path when triggered by a target sound event.
[0239] Target sound events refer to transient acoustic events that have a clear indication of operational status, such as the starting / acceleration roar of a train motor or the braking friction sound before entering a station.
[0240] When the subway train is running smoothly through long tunnels, the mobile terminal is in the first working mode by default. At this time, the terminal continuously runs a lightweight sound event detection algorithm, focusing only on recognizing a few distinctive sound events such as the train motor starting / accelerating and braking.
[0241] Specifically, a finite state machine based on audio event detection can be run. The following features of the audio stream are continuously calculated: (1) Short-Term Energy (STE) and Zero-Crossing Rate (ZCR): used to detect the starting point of sound events. This is because motor starting is usually accompanied by a sharp rise in STE and a specific change pattern of ZCR. (2) Dynamic trajectory of MFCC: Motor sound and braking sound have unique trajectory patterns in the MFCC feature space. (3) Spectral centroid and roll-off point: The spectral centroid of motor acceleration sound will shift to lower frequencies, while the friction sound of brake pads may show an instantaneous increase in the energy of high-frequency components (such as 2-4kHz). Through a lightweight classifier, such as a Support Vector Machine (SVM) or a small-scale Recurrent Neural Network (RNN), based on the feature sequence extracted above, it is determined in real time whether a motor starting, acceleration or braking event has occurred.
[0242] When a target sound event is detected, the system switches from the first working mode to the second working mode.
[0243] Step S217: In the second working mode, collect the current subway environment audio and extract the target acoustic feature vector of the current subway environment audio.
[0244] In the second working mode, the mobile terminal initiates a complete acoustic processing pipeline, recording a high-quality ambient audio segment lasting several seconds (denoted as the current subway ambient audio) and extracting the acoustic feature vector (denoted as the target acoustic feature vector). The terminal uploads the vector and waits.
[0245] Furthermore, after the mobile terminal receives the target location from the server, the mobile terminal returns from the second working mode to the first working mode.
[0246] The subway positioning method provided in this application achieves a balance between high-precision positioning and long battery life of the terminal by enabling high-power computing only at key nodes, thus solving the energy consumption bottleneck problem of mobile terminals in such applications.
[0247] This application provides a subway positioning system.
[0248] The subway positioning system includes a server and a mobile terminal, wherein: The server is used to execute the subway positioning method applied to the server as described in the above embodiments; The mobile terminal is used to execute the subway positioning method applied to the mobile terminal as described in the above embodiments.
[0249] The mobile terminal is not only a sound collector but also a node with basic intelligent processing capabilities. Its core function is to provide high-quality data for the backend positioning service while protecting user privacy and conserving power. The server, deployed in a remote cloud, does not directly collect sound but is responsible for the most complex calculation, decision-making, and storage tasks, and is the core of achieving high-precision positioning.
[0250] The mobile terminal first acquires the target acoustic feature vector of the subway environment audio at the location of the target object, and then sends the target acoustic feature vector to the server. The server matches the target acoustic feature vector with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object. The subway tunnel environment acoustic signature database is constructed based on the first subway characteristic noise collected in the subway tunnel, which includes at least one of wheel-rail noise and wind noise. Then, based on the candidate location set, the target location of the target object is determined and sent to the mobile terminal. The specific execution process can be referred to the above embodiment, and will not be repeated here.
[0251] The subway positioning system provided in this application embodiment constructs a subway tunnel environment acoustic signature database and matches the target acoustic feature vector of the subway environment audio collected by the mobile terminal with the acoustic signature model in the database, thereby achieving continuous and autonomous positioning in environments without traditional signals such as GNSS and base stations. Through this method, continuous, stable, and high-precision positioning is achieved in a closed subway environment without increasing additional subway infrastructure construction costs.
[0252] The following describes the subway positioning device applied to the server provided in this application. The subway positioning device described below and the subway positioning method applied to the server described above can be referred to in correspondence.
[0253] Figure 7 This is one of the structural schematic diagrams of the subway positioning device provided in the embodiments of this application, such as... Figure 7 As shown, the device includes a first acquisition module 710, a feature matching module 720, and a location determination module 730; wherein: The first acquisition module 710 is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located. The feature matching module 720 is used to match the target acoustic feature vector with the subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway feature noise collected in the subway tunnel, and the first subway feature noise includes at least one of wheel-rail noise and wind noise; The location determination module 730 is used to determine the target location of the target object based on the candidate location set.
[0254] The subway positioning device provided in this application obtains the target acoustic feature vector of the subway environment audio where the target object is located through a server; then, it matches the target acoustic feature vector with a subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, the first subway characteristic noise including at least one of wheel-rail noise and wind noise; and then, based on the candidate location set, the target positioning location of the target object is determined. This application embodiment uses the inherent characteristic noise in the subway tunnel environment as an acoustic fingerprint for positioning, without relying on external satellite signals, overcoming the problem of physical layer interruption of positioning caused by satellite signal blockage in the tunnel; at the same time, this method also does not require large-scale deployment of external hardware facilities such as Bluetooth beacons or cellular base stations in the subway tunnel, thereby avoiding high construction and maintenance costs. In addition, compared with radio frequency signals, acoustic features are minimally affected by Faraday shielding of the metal car body and tunnel multipath reflection effects, avoiding the problems of drastic fluctuations in signal strength and frequent jumps in positioning results in traditional solutions, thus improving the positioning accuracy and stability in the closed subway environment. In summary, the embodiments of this application achieve continuous, stable, and high-precision positioning in a closed subway environment without increasing the cost of subway infrastructure construction.
[0255] In one embodiment, the subway positioning device further includes: The third acquisition module is used to acquire sample audio and corresponding geographical location information from multiple collection points in the subway tunnel, and to acquire interference noise samples; the sample audio includes wheel-rail noise and / or wind noise, and the interference noise samples include at least one of human voice, subway broadcast, door opening and closing sound and construction noise; The first extraction module is used to process the sample audio based on the interference noise sample and extract the sample acoustic feature vector from the processed sample audio. The grid partitioning module is used to divide the subway line into multiple geographic grids, and to partition the sample acoustic feature vector into the corresponding geographic grids according to the geographic location information; The voiceprint database training module is used to train the subway tunnel environment voiceprint database based on the segmented sample acoustic feature vectors.
[0256] In one embodiment, the first extraction module includes: The preprocessing unit is used to perform noise reduction and normalization processing on the sample audio based on the interference noise sample to obtain standard audio. The frame segmentation processing unit is used to segment the standard audio into frames to obtain a standard audio frame sequence. The feature extraction unit is used to extract features from the standard audio frame sequence to obtain a time-frequency acoustic feature sequence; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum feature; The feature aggregation unit is used to aggregate and generate the sample acoustic feature vector based on the time-frequency acoustic feature sequence.
[0257] In one embodiment, the preprocessing unit is specifically used for: The sample audio is converted to a time-frequency signal to obtain a sample time-frequency spectrum. Based on the interference noise samples, identify the first random interference noise in the sample audio and identify the first subway characteristic noise in the sample audio; The first random interference noise is filtered, and the target frequency band of the first subway characteristic noise is enhanced to extract the target subway characteristic parameters. The target subway feature parameters and the sample time-frequency graph are input into a preset neural network model to obtain the feature mask output by the preset neural network model. Based on the feature mask, the time-frequency spectrum of the sample is weighted and reconstructed and transformed in the time domain to obtain the processed sample audio. The processed sample audio is then subjected to global level normalization and spectral tilt compensation in sequence to obtain the standard audio.
[0258] In one embodiment, the voiceprint database training module is specifically used for: The acoustic feature vectors of samples from multiple geographic grids belonging to the same metro line are aggregated and trained to generate a line-level acoustic model corresponding to the metro line. The acoustic feature vectors of samples from multiple geographic grids belonging to the same driving section are aggregated and trained to generate a section-level acoustic model corresponding to the driving section. The grid-level acoustic model corresponding to the geographic grid is generated by training based on the acoustic feature vectors of samples divided into the same geographic grid. The subway tunnel environmental acoustic model database includes the line-level acoustic model, the section-level acoustic model, and the grid-level acoustic model.
[0259] In one embodiment, the subway tunnel environmental acoustic fingerprint database includes a line-level acoustic model, a section-level acoustic model, and a grid-level acoustic model. The feature matching module 720 includes: The first matching unit is used to match the target acoustic feature vector with the line-level acoustic model to determine the candidate line set; The second matching unit is used to match the target acoustic feature vector with the interval-level acoustic model corresponding to the candidate line set to determine the candidate interval set; The third matching unit is used to match the acoustic feature vector with the grid-level acoustic model corresponding to the candidate interval set to determine the candidate location set.
[0260] In one embodiment, the location determination module 730 includes: The first acquisition unit is used to acquire the historical location and first timestamp of the target object; The first determining unit is used to determine the theoretically achievable range based on the second timestamp corresponding to the target acoustic feature vector, the historical positioning location, and the preset maximum running speed of the first timestamp. A location filtering unit is used to filter the candidate location set based on the theoretically reachable range to obtain the remaining candidate locations; The second determining unit is used to determine the target location based on the remaining candidate locations.
[0261] In one embodiment, the second determining unit is specifically used for: Obtain the operational constraint information of the subway environment in which the target object is located; the operational constraint information includes at least one of path topology, speed limit rules, and direction of travel; Based on the operational constraint information, the voiceprint matching scores corresponding to the remaining candidate positions are weighted to obtain a comprehensive confidence score. Based on the comprehensive confidence score, the target location is determined from the remaining candidate locations.
[0262] It should be noted that the subway positioning device provided in this application embodiment can implement all the method steps implemented in the above-mentioned subway positioning method embodiment applied to the server, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0263] The following describes the subway positioning device for mobile terminals provided in this application. The subway positioning device described below and the subway positioning method for mobile terminals described above can be referred to in correspondence.
[0264] Figure 8This is a second structural schematic diagram of the subway positioning device provided in the embodiments of this application, as shown below. Figure 8 As shown, the device includes a second acquisition module 810, a feature transmission module 820, and a location receiving module 830; wherein: The second acquisition module 810 is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located; Feature sending module 820 is used to send the target acoustic feature vector to the server; The location receiving module 830 is used to receive the target location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
[0265] The subway positioning device provided in this application embodiment acquires the target acoustic feature vector of the subway environment audio of the target object through a mobile terminal; then, it sends the target acoustic feature vector to a server; and then receives the target positioning location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database. The subway tunnel environment acoustic fingerprint database is constructed based on first subway characteristic noise collected within the subway tunnel, which includes at least one of wheel-rail noise and wind noise. Through this method, the mobile terminal only needs to acquire the acoustic feature vector of the subway environment audio of the target object, without relying on high-precision inertial navigation devices or specific radio signal scanning modules. Environmental data acquisition can be completed using the existing microphone of the mobile terminal, solving the technical deficiency of ordinary mobile terminals that cannot maintain accurate positioning over long distances due to low-precision inertial devices. Simultaneously, the mobile terminal acquires the acoustic feature vector locally and sends it to the server, which performs acoustic fingerprint database matching based on the subway tunnel environment acoustic fingerprint database. Subsequently, the mobile terminal directly receives the high-precision target positioning location. This mode avoids the problem of mobile terminals needing long periods of continuous sampling to converge due to radio frequency multipath effects in enclosed metal carriages. It utilizes acoustic characteristics to achieve rapid response under extremely short acquisition windows. At the same time, it offloads the core matching computing power to the server, ensuring that the terminal device can continuously and stably obtain location updates in the subway carriage without increasing the additional computing burden on the mobile terminal.
[0266] In one embodiment, the second acquisition module 810 includes: The audio acquisition unit is used to collect the current subway environment audio. A noise identification unit is used to identify the second random interference noise, fixed equipment noise, and second subway characteristic noise in the current subway environment audio. An audio processing unit is used to filter the second random interference noise and the fixed equipment noise, and to enhance the target frequency band of the second subway characteristic noise to obtain the target audio. The second extraction unit is used to extract the target acoustic feature vector based on the target audio.
[0267] In one embodiment, the second acquisition module 810 further includes: An event monitoring unit is used to monitor target sound events in the subway environment where the target object is located in the first working mode; The mode switching unit is used to switch from the first working mode to the second working mode when the target sound event is detected; The third extraction unit is used to collect the current subway environment audio in the second working mode and extract the target acoustic feature vector of the current subway environment audio.
[0268] It should be noted that the subway positioning device provided in this application embodiment can implement all the method steps implemented in the above-mentioned subway positioning method embodiment applied to mobile terminals, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0269] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 910, a communications interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communications interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute the subway positioning method provided in the above embodiments.
[0270] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0271] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the subway positioning method provided in the above embodiments.
[0272] In another aspect, this application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the subway positioning method provided in the above embodiments.
[0273] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0274] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0275] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A subway positioning method, characterized in that, Applied to the server side, including: Obtain the target acoustic feature vector of the subway environment audio where the target object is located; The target acoustic feature vector is matched with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic signature database is constructed based on the first subway feature noise collected in the subway tunnel, and the first subway feature noise includes at least one of wheel-rail noise and wind noise; Based on the candidate location set, the target location of the target object is determined.
2. The subway positioning method according to claim 1, characterized in that, Before matching the target acoustic feature vector with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object, the method further includes: Acquire sample audio and corresponding geographical location information from multiple collection points within a subway tunnel, and acquire interference noise samples; the sample audio includes wheel-rail noise and / or wind noise, and the interference noise samples include at least one of human voice, subway broadcasts, door opening and closing sounds, and construction noise; Based on the interference noise samples, the sample audio is processed, and the sample acoustic feature vector is extracted from the processed sample audio. The subway line is divided into multiple geographic grids, and the sample acoustic feature vectors are assigned to the corresponding geographic grids based on the geographic location information. The acoustic signature database for the subway tunnel environment is obtained by training based on the segmented sample acoustic feature vectors.
3. The subway positioning method according to claim 2, characterized in that, The step of processing the sample audio based on the interference noise sample and extracting the sample acoustic feature vector from the processed sample audio includes: Based on the interference noise samples, the sample audio is sequentially subjected to noise reduction and normalization processing to obtain standard audio; The standard audio is segmented into frames to obtain a standard audio frame sequence; Feature extraction is performed on the standard audio frame sequence to obtain a time-frequency acoustic feature sequence; the time-frequency acoustic feature sequence includes at least one of Mel frequency cepstral coefficient feature vector and Mel spectrum feature; Based on the time-frequency acoustic feature sequence, the sample acoustic feature vector is generated by aggregation.
4. The subway positioning method according to claim 3, characterized in that, Based on the interference noise samples, the sample audio is sequentially subjected to noise reduction and normalization processing to obtain standard audio, including: The sample audio is converted to a time-frequency signal to obtain a sample time-frequency spectrum. Based on the interference noise samples, identify the first random interference noise in the sample audio and identify the first subway characteristic noise in the sample audio; The first random interference noise is filtered, and the target frequency band of the first subway characteristic noise is enhanced to extract the target subway characteristic parameters. The target subway feature parameters and the sample time-frequency graph are input into a preset neural network model to obtain the feature mask output by the preset neural network model. Based on the feature mask, the time-frequency spectrum of the sample is weighted and reconstructed and transformed in the time domain to obtain the processed sample audio. The processed sample audio is then subjected to global level normalization and spectral tilt compensation in sequence to obtain the standard audio.
5. The subway positioning method according to claim 2, characterized in that, The process of training based on the partitioned sample acoustic feature vectors to obtain the subway tunnel environmental acoustic signature database includes: The acoustic feature vectors of samples from multiple geographic grids belonging to the same metro line are aggregated and trained to generate a line-level acoustic model corresponding to the metro line. The acoustic feature vectors of samples from multiple geographic grids belonging to the same driving section are aggregated and trained to generate a section-level acoustic model corresponding to the driving section. The grid-level acoustic model corresponding to the geographic grid is generated by training based on the acoustic feature vectors of samples divided into the same geographic grid. The subway tunnel environmental acoustic model database includes the line-level acoustic model, the section-level acoustic model, and the grid-level acoustic model.
6. The subway positioning method according to claim 1, characterized in that, The subway tunnel environment acoustic signature database includes a line-level acoustic model, a section-level acoustic model, and a grid-level acoustic model. Matching the target acoustic feature vector with the subway tunnel environment acoustic signature database to determine the candidate location set corresponding to the target object includes: The target acoustic feature vector is matched with the line-level acoustic model to determine the candidate line set; The target acoustic feature vector is matched with the interval-level acoustic model corresponding to the candidate line set to determine the candidate interval set; The acoustic feature vectors are matched with the grid-level acoustic models corresponding to the candidate interval set to determine the candidate location set.
7. The subway positioning method according to any one of claims 1 to 6, characterized in that, Determining the target location of the target object based on the candidate location set includes: Obtain the historical location and first timestamp of the target object; Based on the second timestamp corresponding to the target acoustic feature vector, the historical positioning location, and the preset maximum running speed of the first timestamp, the theoretically achievable range is determined. Based on the theoretically reachable range, the candidate location set is filtered to obtain the remaining candidate locations; Based on the remaining candidate locations, the target location is determined.
8. The subway positioning method according to claim 7, characterized in that, Determining the target location based on the remaining candidate locations includes: Obtain the operational constraint information of the subway environment in which the target object is located; the operational constraint information includes at least one of path topology, speed limit rules, and direction of travel; Based on the operational constraint information, the voiceprint matching scores corresponding to the remaining candidate positions are weighted to obtain a comprehensive confidence score. Based on the comprehensive confidence score, the target location is determined from the remaining candidate locations.
9. A subway positioning method, characterized in that, Applied to mobile terminals, including: Obtain the target acoustic feature vector of the subway environment audio at the target object location; Send the target acoustic feature vector to the server; The server receives the target location returned by the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database; wherein the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
10. The subway positioning method according to claim 9, characterized in that, The acquisition of the target acoustic feature vector of the subway environment audio of the target object includes: Collect audio from the current subway environment; Identify the second random interference noise, fixed equipment noise, and second subway characteristic noise in the current subway environment audio; The second random interference noise and the fixed equipment noise are filtered, and the target frequency band of the second subway characteristic noise is enhanced to obtain the target audio. Based on the target audio, the target acoustic feature vector is extracted.
11. The subway positioning method according to claim 9, characterized in that, The acquisition of the target acoustic feature vector of the subway environment audio of the target object includes: In the first working mode, target sound events are monitored in the subway environment where the target object is located; When the target sound event is detected, the system switches from the first working mode to the second working mode. In the second working mode, the current subway environment audio is collected, and the target acoustic feature vector of the current subway environment audio is extracted.
12. A subway positioning device, characterized in that, include: The first acquisition module is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located; The feature matching module is used to match the target acoustic feature vector with the subway tunnel environment acoustic fingerprint database to determine the candidate location set corresponding to the target object; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise; The location determination module is used to determine the target location of the target object based on the candidate location set.
13. A subway positioning device, characterized in that, include: The second acquisition module is used to acquire the target acoustic feature vector of the subway environment audio where the target object is located; The feature sending module is used to send the target acoustic feature vector to the server; The location receiving module is used to receive the target location returned by the server based on the target acoustic feature vector and the subway tunnel environment acoustic fingerprint database; wherein, the subway tunnel environment acoustic fingerprint database is constructed based on the first subway characteristic noise collected in the subway tunnel, and the first subway characteristic noise includes at least one of wheel-rail noise and wind noise.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the subway positioning method as described in any one of claims 1 to 8, or the subway positioning method as described in any one of claims 9 to 11.
15. A subway positioning system, characterized in that, This includes server-side and mobile terminals, among which: The server is used to execute the subway positioning method as described in any one of claims 1 to 8; The mobile terminal is used to execute the subway positioning method as described in any one of claims 9 to 11.
16. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the subway positioning method as described in any one of claims 1 to 8, or the subway positioning method as described in any one of claims 9 to 11.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the subway positioning method as described in any one of claims 1 to 8, or the subway positioning method as described in any one of claims 9 to 11.