Train positioning matching method based on voice recognition technology

Through a train positioning method based on speech recognition technology, voice conversion and urban rail transit context are used to screen correct information, solving the problem of inaccurate train positioning when the interlocking host fails, improving positioning efficiency and accuracy, and reducing reliance on human judgment.

CN120636403APending Publication Date: 2025-09-12浙江众合科技股份有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510563347.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In urban rail transit, when the interlocking host fails, the dispatcher cannot accurately obtain the train location, resulting in long manual confirmation time and prone to errors, which may cause safety accidents.

Method used

A train positioning method based on speech recognition technology is adopted. By obtaining the position report audio, it is converted into text information using speech recognition technology. Combined with the contextual semantics of urban rail transit and preset keyword requirements, correct and incorrect information is screened, and accurate or fuzzy generation strategies are selected to generate train positioning information.

Benefits of technology

It improves the accuracy and efficiency of train positioning, reduces reliance on human judgment, avoids mispositioning, and ensures fast and accurate decision-making by the dispatching center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636403A_ABST
    Figure CN120636403A_ABST
Patent Text Reader

Abstract

The invention discloses a train positioning matching method based on a voice recognition technology, and relates to the technical field of urban rail transit, and the method comprises the following steps: obtaining a position report audio, and obtaining a first conversion result based on the voice recognition technology according to the position report audio; obtaining a second conversion result according to the first conversion result and a context semantic correlation matching result; screening report correct information and report error information according to a matching result of the second conversion result and a preset keyword requirement; and according to the report correct information and the report error information, selecting to execute an accurate generation strategy or a fuzzy generation strategy so as to generate train positioning information. The method has the beneficial effects that the overall train positioning confirmation time cost can be shortened, and the train positioning accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of urban rail transit technology, and in particular to a train positioning and matching method based on speech recognition technology. Background Art

[0002] With the development of urban rail transit, in order to further shorten the impact time of equipment failure on rail transit operations, higher requirements are placed on the efficiency of central dispatching and handling command when equipment failure occurs. At present, when the interlocking host fails completely, the dispatching interface cannot obtain information about the train and trackside equipment, and the dispatcher cannot grasp the position of the trains operating on the main line. It is necessary to communicate with the train drivers one by one by radio in order to confirm the train position; due to the large number of trains in urban rail transit lines, the entire train position confirmation process takes a long time. At the same time, the train position confirmation relies entirely on human judgment. Once an information verification error occurs, it may cause a major safety accident. Therefore, there is a need for a technology that can quickly assist dispatching to complete the positioning of trains on the entire line when the interlocking host fails completely, which can not only shorten the overall train positioning confirmation time cost, but also improve the accuracy of train positioning. Summary of the Invention

[0003] This application addresses the problem in the prior art that train positioning relies on human judgment and has low accuracy. It provides a train positioning matching method based on speech recognition technology, which obtains text conversion results based on the actual context of the position report audio, further judges the correctness of the content in the text conversion result based on preset keyword requirements, and performs accurate generation and fuzzy generation of train positioning information respectively according to whether the report content is completely correct. While providing a train positioning reference for central dispatching, it can also remind staff to pay attention when there is erroneous reporting information, without relying on the monitoring results of other monitoring equipment, thereby improving the efficiency and accuracy of train positioning.

[0004] In order to achieve the above-mentioned technical objectives, a technical solution provided in this application is a train positioning matching method based on speech recognition technology, comprising the following steps: obtaining position report audio, and obtaining a first conversion result based on the position report audio based on speech recognition technology; obtaining a second conversion result based on the first conversion result and the context semantic relevance matching result; screening correct reporting information and incorrect reporting information based on the matching result of the second conversion result and the preset keyword requirements; and selecting to execute an accurate generation strategy or a fuzzy generation strategy based on the correct reporting information and the incorrect reporting information to generate train positioning information.

[0005] Furthermore, obtaining the second conversion result based on the first conversion result and the context semantic relevance matching result includes: constructing context semantic relevance based on the matching between text and semantics in the urban rail transit context; and obtaining the second conversion result based on the first conversion result and the context semantic relevance.

[0006] Furthermore, obtaining the second conversion result based on the first conversion result and the context semantic relevance matching result includes: constructing context semantic relevance based on the matching of text and semantics in each urban rail transit context; retrieving the corresponding context semantic relevance based on the current urban rail transit, and obtaining the second conversion result based on the first conversion result and the corresponding context semantic relevance.

[0007] Furthermore, the screening and reporting of correct information and reporting of error information based on the matching results of the second conversion result and the preset keyword requirements includes: obtaining the preset keyword requirements according to the reporting template; extracting key information according to the second conversion result; if the key information matches the preset keyword requirements, it is reported as correct information, and if the key information does not match the preset keyword requirements, it is reported as error information.

[0008] Furthermore, the key information includes train number, vehicle body number, station, section, up and down lines, storage line and kilometer mark.

[0009] Furthermore, the screening and reporting of correct information and reporting of error information based on the matching results of the second conversion result and the preset keyword requirements includes: obtaining the preset keyword requirements based on the reporting template and the keyword association relationship; if the key information in the second conversion result matches the preset keyword requirements, it is reported as correct information; if the key information in the second conversion result does not match the preset keyword requirements, it is reported as error information.

[0010] Furthermore, the selection of executing an accurate generation strategy or a fuzzy generation strategy to generate train positioning information based on the reported correct information and the reported incorrect information includes: if only correct information is reported, executing the accurate generation strategy to generate train positioning information; if there is reported incorrect information, executing the fuzzy generation strategy to generate train positioning information.

[0011] Furthermore, the execution of the accurate generation strategy to generate train positioning information includes: generating a train number window in the corresponding section based on the location information in the reported correct information; generating train number and vehicle body information in the train number window based on the train information in the reported correct information; and generating train positioning information using the train number window, train number and vehicle body information.

[0012] Furthermore, the execution of the fuzzy generation strategy to generate train positioning information includes: judging whether to execute fuzzy generation based on whether there is necessary positioning information in the reported error information; if fuzzy generation is executed, generating train positioning information based on the reported correct information; if fuzzy generation is not executed, generating positioning failure information based on the reported error information.

[0013] Furthermore, the necessary positioning information includes at least the train number, vehicle body number, kilometer mark, and up and down lines.

[0014] The beneficial effects of this application are as follows: based on the audio position report, the audio information is converted into text information based on speech recognition technology to obtain a first conversion result. Since the semantics of different texts may differ, in order to further ensure the accuracy of recognition, the corresponding semantics are matched according to the current urban rail transit context of the train to obtain a second conversion result. According to the second conversion result and the preset keyword requirements, matching is performed to screen out the correct information and error information in the audio position report, thereby avoiding the train's mispositioning caused by the driver's reporting error, improving the accuracy of train positioning, and without relying on the normal operation of other monitoring equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flow chart of a train positioning and matching method based on speech recognition technology in this application.

[0016] Figure 2 This is a schematic diagram of a station type in an embodiment of the present application. Figure 1 .

[0017] Figure 3 This is a schematic diagram of a station type in an embodiment of the present application. Figure 2 . DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this application more clear, the application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific implementation method described here is only an optimal embodiment of this application, which is only used to explain this application and does not limit the scope of protection of this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0019] like Figure 1 As shown in the first embodiment of the present application, a train positioning and matching method based on speech recognition technology includes the following steps: Acquire a location report audio, and obtain a first conversion result based on the location report audio based on speech recognition technology; Obtaining a second conversion result based on a matching result between the first conversion result and the contextual semantic relevance; Screening and reporting correct information and reporting error information based on the matching result between the second conversion result and the preset keyword requirement; According to the reported correct information and the reported incorrect information, an accurate generation strategy or a fuzzy generation strategy is selected to generate train positioning information.

[0020] When an anomaly occurs, in order to ensure that the dispatch center can quickly analyze key information, train drivers usually need to report according to a fixed template. Therefore, in this embodiment, when the driver reports the train's location in a voice according to the established template, a position report audio is obtained. Based on the position report audio, the audio information is converted into text information based on speech recognition technology to obtain a first conversion result. Since the semantics of different texts may differ, in order to further ensure the accuracy of recognition, the corresponding semantics are matched according to the current urban rail transit context of the train to obtain a second conversion result. Based on the second conversion result and the preset keyword requirements, a match is performed to screen out the correct information and error information in the position report audio, thereby avoiding the train's mispositioning caused by the driver's reporting error, improving the accuracy of train positioning, and without relying on the normal operation of other monitoring equipment.

[0021] After obtaining the location report audio, preprocessing is performed on the audio. This preprocessing includes at least filtering out operational noise. In this embodiment, spectral subtraction is used to preprocess the location report audio, removing operational noise from the audio and ensuring more accurate subsequent speech recognition results. In other cases, preprocessing can also be performed using deep learning speech denoising, sub-control algorithms, and other methods.

[0022] Speech recognition technology converts the vocabulary content in human speech into computer-readable input. A neural network model can be used to learn the feature matching between audio and text, and the text can be matched according to the features of the location report audio, thereby realizing the conversion of audio to text. In this embodiment, the Mel-frequency cepstral coefficients are used to extract the acoustic features in the location report audio, and the acoustic features are converted into phonemes using an acoustic model. The phonemes are converted into text information output using a language model to realize the conversion of audio to text. Among them, the acoustic model can adopt the DNN (deep neural network) + CTC (connectionist temporal classification) architecture, and the language model can adopt the N-gram language model.

[0023] Acquiring a second conversion result according to the first conversion result and the context semantic relevance matching result includes: Construct contextual semantic relevance based on the matching between text and semantics in the urban rail transit context; A second conversion result is obtained based on the semantic relevance between the first conversion result and the context.

[0024] The first conversion result is a conventional text conversion from speech audio recognition. However, urban rail transit contains a large number of specialized terminology. If conventional semantics are used for recognition, it is easy to fail to correctly understand the report content, such as incorrect sentence segmentation and integration, which may cause errors in subsequent recognition. Therefore, this embodiment pre-establishes contextual semantic relevance based on the matching of text and semantics in actual urban rail transit. This contextual semantic relevance serves as the basis for text semantic recognition, reflecting the specialized terminology in urban rail transit. This ensures that the recognition results are consistent with the actual urban rail transit reports and improves the accuracy of train positioning.

[0025] In other cases, obtaining the second conversion result according to the first conversion result and the context semantic relevance matching result includes: Construct contextual semantic relevance based on the matching between text and semantics in each urban rail transit context; The corresponding contextual semantic relevance is retrieved according to the current urban rail transit, and the second conversion result is obtained according to the first conversion result and the corresponding contextual semantic relevance.

[0026] In this case, considering that different urban rail transit systems may have different descriptions, we construct contextual semantic relevance for each context based on the matching between text and semantics in each urban rail transit context. This allows us to retrieve the corresponding contextual semantic relevance based on the current urban rail transit and perform semantic matching based on the text conversion results, further improving recognition accuracy and avoiding recognition discrepancies caused by regional differences.

[0027] In some embodiments, screening and reporting correct information and reporting error information based on the matching result between the second conversion result and the preset keyword requirement includes: Obtain preset keyword requirements based on the report template; extracting key information according to the second conversion result; If the key information matches the preset keyword requirements, it is reported as correct information. If the key information does not match the preset keyword requirements, it is reported as incorrect information.

[0028] Since the train driver reports according to the established template, preset keyword requirements are set according to the reporting template. When the key information in the second conversion result matches any preset keyword requirement, it is considered to be correct reporting information. If the key information in the second conversion result does not match any preset keyword requirement, it is considered to be incorrect reporting information, thereby screening out the incorrect information in the train driver's report to improve the accuracy of subsequent positioning.

[0029] In this embodiment, key information includes train information and location information. Train information includes train number and vehicle body number. Location information includes station, platform section, up and down lines, storage line and kilometer mark. The preset keyword requirements are set according to the provisions in the established template, such as the number of characters in the train number and the number of characters in the vehicle body number. In other cases, the preset keyword requirements can also be set according to the actual situation. For example, if the train number only exists between 201 and 210, the preset keyword requirement is that the train number exists between 201 and 210. If the train number position in the second conversion result is 211, it is considered that it does not match the preset keyword requirement and is reported as an error message.

[0030] In other embodiments, screening and reporting correct information and reporting error information based on the matching result between the second conversion result and the preset keyword requirement includes: Obtain preset keyword requirements and keyword order based on the report template; Matching the second conversion result with the preset keyword requirements in sequence according to the keyword order; If the key information in the second conversion result matches the preset keyword requirement, it is reported as correct information. If the key information in the second conversion result does not match the preset keyword requirement, it is reported as incorrect information.

[0031] In this embodiment, the key information must not only match the preset keyword requirements, but also match the keyword order, so as to avoid misidentification when the train driver reports omissions. For example, when the train driver reports the train number, he should report: train number -201, but the train driver missed reporting it as: 201, omitting the train number. At this time, when matching according to the keyword order, it can be known that the train driver reported the train number 201, not the vehicle body number or other information, thereby improving the accuracy of audio recognition, and providing a certain amount of fault tolerance for the train driver, thereby improving the efficiency and accuracy of train positioning.

[0032] In some further embodiments, screening and reporting correct information and reporting error information based on the matching result between the second conversion result and the preset keyword requirement includes: Obtain preset keyword requirements based on the report template and keyword associations; If the key information in the second conversion result matches the preset keyword requirement, it is reported as correct information. If the key information in the second conversion result does not match the preset keyword requirement, it is reported as incorrect information.

[0033] Keyword association relationships include conflict and association relationships between individual keyword information. Keyword association relationships are obtained based on the actual route, such as the association between kilometer marker information and interval information, and the conflict relationship between interval and platform information. When the kilometer marker information and interval information in the key information both correspond, both key information are considered to be correctly reported. When the kilometer marker information and interval information do not correspond, the interval information is included in the report as an error. When both interval and platform appear in the key information, both key information are considered to be incorrectly reported. Further screening of correct and incorrect information in the second conversion result improves the accuracy of train positioning information generation.

[0034] Selecting to execute an accurate generation strategy or a fuzzy generation strategy based on reporting correct information and reporting incorrect information to generate train positioning information includes: If only correct information is reported, the accurate generation strategy is executed to generate train positioning information; If there is reported error information, the fuzzy generation strategy is executed to generate train positioning information.

[0035] In this embodiment, the accurate generation strategy or the fuzzy generation strategy is selectively executed based on whether there is reported error information, so that when there is reported error information, positioning based on the error information can be avoided. While providing a reference for central scheduling, the existence of reported error information is prompted, which facilitates staff to conduct secondary confirmation.

[0036] The execution of an accurate generation strategy to generate train positioning information includes: Generate a train window in the corresponding section based on the position information in the correct information reported; Generate train number and vehicle body information in the train number window according to the train information in the correct information reported; Generate train location information based on train window, train number and car body information.

[0037] Location information includes the station, platform section, up and down lines, storage lanes, and kilometer markers. Train information includes the train number and car body number.

[0038] The fuzzy generation strategy is implemented to generate train location information including: Determine whether to execute the fuzzy generation strategy based on whether there is necessary positioning information in the reported error information; If the fuzzy generation strategy is implemented, the train location information is generated based on the reported correct information; If the fuzzy generation strategy is not executed, the positioning failure information is generated based on the reported error information.

[0039] In this embodiment, if the error report contains essential positioning information, the fuzzy generation strategy is not executed. If the error report does not contain essential positioning information, the fuzzy generation strategy is executed to ensure the accuracy of the positioning information provided to the central dispatcher. When the fuzzy generation strategy is executed, a train number window is generated in the corresponding section based on the location information in the correct report. The train number and vehicle body information are generated in the train number window based on the train information in the correct report. The train positioning information is generated using the train number window, train number, and vehicle body information. If the fuzzy generation strategy is not executed, a positioning failure message is generated based on the error report and uploaded to the central dispatcher for easy processing by staff.

[0040] Necessary positioning information includes at least the train number, vehicle body number, kilometer marker, and uplink and downlink directions. If any necessary positioning information is included in the error report, the fuzzy generation strategy is not executed. If no necessary positioning information is included in the error report, a train window is generated in the corresponding section based on the uplink and downlink directions and kilometer markers, and train number and vehicle body information are generated in the train window based on the train number and vehicle body number.

[0041] In other embodiments, when the error report contains conflicting key information, a positioning failure message is generated based on the error report. For example, if the error report contains both "platform" and "section," since the train cannot be at both the platform and the section at the same time, conflicting key information is considered to exist.

[0042] In some cases, in order to reduce information conflicts and energy consumption, the position reporting audio is acquired and recognized only when the central dispatch fails due to an interlocking host failure.

[0043] As the second embodiment of the present application, a specific embodiment is provided for illustration. Define right as the uplink direction of the line, left as the downlink direction of the line, according to the station type (such as Figure 2 and Figure 3 (As shown, there were two trains on the line: Train 201, Car 001, and Train 203, Car 005. Due to a fault in the interlocking host, the central dispatcher activated the voice recognition function, prompting the driver to report the train's position to the central dispatcher. The driver of Train 201 reported that Train 001 had stopped at the designated platform on the downlink section of Station X. The driver of Train 202 reported that Train 203, Car 005, had stopped at K2 + 500 meters on the uplink section between Stations Y and Z.

[0044] At this time, the first position report audio obtained is: train number-201, vehicle body number-001, station-X station, platform / section / storage line-platform, up / down-down; the second position report audio obtained is: train number-203, vehicle body number-005, station-Y station to Z station, platform / section / storage line-section, up / down-up, kilometer mark-K2+500.

[0045] Based on speech recognition technology, a first conversion result corresponding to the first location report audio and a first conversion result corresponding to the second location report audio are obtained. Based on the first conversion result corresponding to the first location report audio and the context semantic relevance matching result, a second conversion result corresponding to the first location report audio is obtained. Based on the first conversion result corresponding to the second location report audio and the context semantic relevance matching result, a second conversion result corresponding to the second location report audio is obtained.

[0046] Screening of correct and incorrect reporting information is performed based on preset keyword requirements: (1) The second conversion result corresponding to the first position report audio: The key information all meets the preset keyword requirements and is reported correctly; (2) The second conversion result corresponding to the second position report audio: K2+500 is between Station X and Station Y, but not between Station Y and Station Z. This does not conform to the actual situation of the line, that is, it does not meet the preset keyword requirements and is reported as an error message.

[0047] According to the correct information reported and the wrong information reported, the accurate generation strategy or the fuzzy generation strategy is selected to generate the train positioning information: (1) The second conversion result corresponding to the first position report audio: Only when correct information is reported and an accurate generation strategy is executed, a train number window is generated in the down platform section of Station X in the interface, and the train number and vehicle body number 201001 are generated in the train number window.

[0048] (2) The second conversion result corresponding to the second position report audio: There is reporting error information, and the fuzzy generation strategy is executed. Since the reporting error information does not include the kilometer mark, a train window is generated in the section corresponding to K2+500 meters. The train number and vehicle body number 203005 are generated in the train window, and the dispatcher is prompted to generate the train window position fuzzily.

[0049] The specific implementation method described above is a preferred implementation method of the train positioning and matching method based on voice recognition technology in this application, and it does not limit the specific implementation scope of this application. The scope of this application includes but is not limited to this specific implementation method. Any equivalent changes made in accordance with the shape and structure of this application are within the scope of protection of this application.

Claims

1. A train positioning and matching method based on speech recognition technology, characterized in that: The steps include: Acquire a location report audio, and obtain a first conversion result based on the location report audio based on speech recognition technology; Obtaining a second conversion result based on a matching result between the first conversion result and the contextual semantic relevance; Screening and reporting correct information and reporting error information based on the matching result between the second conversion result and the preset keyword requirement; According to the reported correct information and the reported incorrect information, an accurate generation strategy or a fuzzy generation strategy is selected to generate train positioning information.

2. The train positioning and matching method based on speech recognition technology according to claim 1, characterized in that: Obtaining the second conversion result according to the first conversion result and the context semantic relevance matching result includes: Construct contextual semantic relevance based on the matching between text and semantics in the urban rail transit context; A second conversion result is obtained based on the semantic relevance between the first conversion result and the context.

3. The train positioning and matching method based on speech recognition technology according to claim 1, characterized in that: Obtaining the second conversion result according to the first conversion result and the context semantic relevance matching result includes: Construct contextual semantic relevance based on the matching between text and semantics in each urban rail transit context; The corresponding contextual semantic relevance is retrieved according to the current urban rail transit, and the second conversion result is obtained according to the first conversion result and the corresponding contextual semantic relevance.

4. The train positioning and matching method based on speech recognition technology according to claim 1, characterized in that: The screening and reporting of correct information and reporting of error information according to the matching result between the second conversion result and the preset keyword requirement includes: Obtain preset keyword requirements based on the report template; extracting key information according to the second conversion result; If the key information matches the preset keyword requirements, it is reported as correct information. If the key information does not match the preset keyword requirements, it is reported as incorrect information.

5. The train positioning and matching method based on speech recognition technology according to claim 4, characterized in that: The key information includes train number, vehicle body number, station, section, up and down lines, storage line and kilometer mark.

6. The train positioning and matching method based on speech recognition technology according to claim 1, characterized in that: The screening and reporting of correct information and reporting of error information according to the matching result between the second conversion result and the preset keyword requirement includes: Obtain preset keyword requirements based on the report template and keyword associations; If the key information in the second conversion result matches the preset keyword requirement, it is reported as correct information. If the key information in the second conversion result does not match the preset keyword requirement, it is reported as incorrect information.

7. The train positioning and matching method based on speech recognition technology according to claim 1, characterized in that: The selecting and executing the accurate generation strategy or the fuzzy generation strategy to generate the train positioning information according to the reported correct information and the reported incorrect information includes: If only correct information is reported, the accurate generation strategy is executed to generate train positioning information; If there is reported error information, the fuzzy generation strategy is executed to generate train positioning information.

8. The train positioning and matching method based on speech recognition technology according to claim 7, characterized in that: The executing of the accurate generation strategy to generate train positioning information includes: Generate a train window in the corresponding section based on the position information in the correct information reported; Generate train number and vehicle body information in the train number window according to the train information in the correct information reported; Generate train location information based on train window, train number and car body information.

9. The train positioning and matching method based on speech recognition technology according to claim 7, characterized in that: The executing of the fuzzy generation strategy to generate train positioning information includes: Determine whether to perform fuzzy generation based on whether there is necessary positioning information in the reported error information; If fuzzy generation is performed, train positioning information is generated based on the reported correct information; If fuzzy generation is not performed, a positioning failure message is generated based on the reported error message.

10. The train positioning and matching method based on speech recognition technology according to claim 9, characterized in that: The necessary positioning information at least includes the train number, vehicle body number, kilometer mark, and up and down lines.