Over-limit event monitoring method based on multi-source real-time data, medium and equipment

By establishing an over-limit event database and comparing multi-source data in real time, the problem of the existing technology being unable to timely identify potential risks during flight has been solved, and timely identification and early warning of over-limit events have been achieved, thereby improving flight safety.

CN120708442AActive Publication Date: 2025-09-26HANGKE TECH DEV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511212553.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies are unable to promptly identify potential risks during flight and fail to fully integrate multiple data resources for real-time risk identification, making it difficult to prevent over-limit incidents.

Method used

Through a method based on multi-source real-time data, an over-limit event database is established. Flight operation information, real-time cabin voice information and QAR information are used to generate the prompt start time, positioning start time and positioning end time of the event to be monitored. The cockpit voice data is compared with the standard voice call information in real time to determine the over-limit event.

Benefits of technology

It enables timely detection of over-limit events during flight, improves the timeliness of flight risk warnings, and reduces the possibility of flight accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708442A_ABST
    Figure CN120708442A_ABST
Patent Text Reader

Abstract

The invention relates to the field of overrun event monitoring, in particular to an overrun event monitoring method based on multi-source real-time data, a medium and equipment. Through fusion of QAR data, flight operation information and SOP, an over-limit event database is established, and prompt and positioning node information is configured for each event to be monitored. The method comprises the following steps: acquiring QAR information of an aircraft in real time, and generating prompt starting time t1, positioning starting time t2 and positioning ending time t3 of an event; and according to a comparison result of the cabin sound information and the SOP standard voice calling information, whether the event exceeds the limit or not is judged in real time. According to the method, each overrun event can be monitored in real time, and the calculation efficiency and the early warning timeliness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of over-limit event monitoring, and in particular to an over-limit event monitoring method, medium and device based on multi-source real-time data. Background Art

[0002] Cockpit voice and flight data are two core data of aircraft, and also important information sources for civil aviation unsafe event investigation, incident investigation and flight quality monitoring. The ex-post analysis application based on cockpit voice and flight data has been relatively mature, but it cannot timely detect potential risks during flight. Moreover, in flight safety monitoring, multiple data resources have not been fully integrated and combined with standard operating procedures (SOP) for real-time risk identification, resulting in difficulty in timely detecting and preventing over-limit events during flight. The over-limit event of cabin voice refers to the output result of checking and judging the compliance of the pilot's standard call according to the standard operating procedure (SOP). As a result, it is impossible to timely determine the possible safety risks of the aircraft during flight, and thus the safety of the aircraft is difficult to be better guaranteed. Summary of the Invention

[0003] For one of the above technical problems, the technical solution adopted by the present invention is as follows: According to one aspect of the present invention, there is provided an over-limit event monitoring method based on multi-source real-time data, and the method includes the following steps: Match the flight monitoring information obtained in the current data collection period with the prompt node information and positioning node information corresponding to each to-be-monitored event in the over-limit event database, so as to generate a prompt start time t1, a positioning start time t2 and a positioning end time t3 corresponding to each to-be-monitored event; t1 < t2 < t3; the to-be-monitored event is a standard operation task including a call task in the SOP corresponding to the current flight; the prompt node information and the positioning node information are respectively time generation strategy information formed by cabin voice information and / or QAR information generated when the preset operation corresponding to the to-be-monitored event in the SOP is executed; the flight monitoring information includes one or any combination of flight operation information, real-time cabin voice information and real-time QAR information; If there is a new complete and valid speech recognition result in the real-time cabin voice information collected in the current data collection period, then traverse t1, t2 and t3 corresponding to each to-be-monitored event; If the end time T of the last complete sentence in the speech recognition results of all valid voice channels obtained in the current data collection period satisfies the condition T ∈ (t1, t2] for any to-be-monitored event corresponding to t1 and t2, then generate the standard voice call information required in the SOP corresponding to the to-be-monitored event; If T and t2 and t3 corresponding to any event to be monitored meet the condition T∈(t2, t3], then each complete speech recognition information obtained in the time window [t2, T] is compared with the standard voice call information required by the SOP corresponding to the event to be monitored; If T and t3 corresponding to any event to be monitored meet the condition T>t3, and if all voice recognition information in all voice comparison results of the event to be monitored within the time window [t2, t3] is different from the standard voice announcement information required in the SOP corresponding to the event to be monitored, then the event to be monitored is determined to be an over-limit event of the flight.

[0004] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned over-limit event monitoring method based on multi-source real-time data.

[0005] According to a third aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method for monitoring over-limit events based on multi-source real-time data is implemented.

[0006] The present invention has at least one of the following beneficial effects: The present invention integrates multiple data sources, including QAR data, flight operation information, and SOPs, to establish a pre-established out-of-limit event database for the current flight (i.e., aircraft). This database includes every event to be monitored (i.e., an out-of-limit event) throughout the flight's operation. Corresponding prompt node information and positioning node information are configured for each event. During the actual flight operation, the prompt start time t1, positioning start time t2, and positioning end time t3 corresponding to each event to be monitored are generated in real time based on the aircraft's real-time QAR information acquired during each data collection cycle. The out-of-limit event monitoring execution period corresponding to each event to be monitored is then formed from t1, t2, and t3. A final determination of whether the event to be monitored is out of limit is then generated based on the comparison results between cockpit voice data acquired in real time during the corresponding determination period and the standard voice announcement information required by the SOP. The present invention enables the real-time comparison of QAR data segments downloaded by the onboard system with the out-of-limit event database to enable the real-time activation of different out-of-limit determination periods corresponding to each event to be monitored, and initiate the corresponding out-of-limit monitoring task. In this way, the various over-limit events that are very complex and have a sequence relationship during the flight process can be decoupled, and over-limit monitoring can be performed separately in real time, thereby improving computing efficiency and timely outputting the results of whether there is an over-limit.

[0007] In addition, the present invention realizes the real-time reception and processing of multi-source data of aircraft. Compared with the traditional post-analysis mode, it can timely detect over-limit events during the flight, greatly improving the timeliness of early warning of flight risks, buying precious time for taking emergency measures, and effectively reducing the possibility of flight accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0009] Figure 1 The present invention provides a flowchart of a method for monitoring over-limit events based on multi-source real-time data in an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of the principle of over-limit event identification and processing provided by an embodiment of the present invention.

[0011] Figure 3 This is a flow chart of a method for obtaining an effective voice channel provided by another embodiment of the present invention.

[0012] Figure 4 A schematic structural diagram of an electronic device provided in another embodiment of the present invention. DETAILED DESCRIPTION

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0014] As a possible embodiment of the present invention, Figure 1 As shown, a method for monitoring over-limit events based on multi-source real-time data is provided, the method comprising the following steps: S100: Match the flight monitoring information obtained in the current data collection period with the prompt node information and positioning node information corresponding to each to-be-monitored event in the over-limit event database to generate the corresponding prompt start time t1, positioning start time t2, and positioning end time t3 for each to-be-monitored event. t1 < t2 < t3; the to-be-monitored events are the standard operation tasks including the shouting tasks in the SOP corresponding to the current flight; the prompt node information and the positioning node information are respectively the time generation strategy information formed by the cabin voice information and / or QAR information (in the aviation field, QAR information is a kind of data for recording flights in real time) generated when the preset operations corresponding to the to-be-monitored events are executed; the flight monitoring information includes one or any combination of flight operation information, real-time cabin voice information, and real-time QAR information. The QAR information in this embodiment is the engineering value after QAR decoding.

[0015] In this embodiment, each to-be-monitored event is actually each necessary standard shouting event performed by the pilot during the flight according to the standard operation procedure (SOP). Some of the to-be-monitored events are shown in Table 1 below: Table 1

[0016] Each row in Table 1 is a to-be-monitored event. Among them, the standard shouting content - PF / L means that this standard shouting content needs to be shouted by the pilot-in-command (captain), and the standard shouting content - PM / R means that this standard shouting content needs to be shouted by the co-pilot.

[0017] Specifically, in this step, obtain the corresponding determination information by receiving the QAR data segment and the cockpit voice data stream transmitted in real time by the airborne system. The specific data collection period can be flexibly set according to the actual scenario. For example, it can be 1 second, and thus the monitoring frequency of over-limit events can be flexibly adjusted.

[0018] Before this step is executed, it is necessary to first establish an over-limit event database for the corresponding flight of this aircraft. Specifically, the over-limit event database can be established according to flight operation information, the text information identified from cabin voice information, SOP, and the historical QAR engineering value after decoding. Specifically, establishing the over-limit event database is actually establishing the corresponding prompt node information and positioning node information for each to-be-monitored event. Among them, the prompt node information is used to locate and generate the pre-reminder time corresponding to the to-be-monitored event, that is, t1. The positioning node information is used to locate and generate the shouting content extraction time corresponding to the to-be-monitored event, that is, t2 and t3.

[0019] Both the prompt node information and the positioning node information are generated based on the inherent logic between related information during actual flight, creating a time-based logic strategy. For example, the generation time of a QAR message can be set to t1, t2, or t3. Alternatively, the QAR information used for positioning at t2 and t3 can be different, or even voice-based. Each t1, t2, and t3 can be generated with a corresponding time offset based on the QAR information.

[0020] Specifically, the QAR retrieval time (t1, t2, and t3) corresponding to each event to be monitored in S100 can be obtained according to the following steps: S101: Matching flight monitoring information acquired during the current data collection period with time-limited generation strategy information of any prompt node information corresponding to the flight in the overrun event database.

[0021] S102: If the currently collected flight monitoring information satisfies the time generation strategy information limited by any prompt node information, then t1 of the event to be monitored corresponding to the prompt node information is generated.

[0022] S103: Matching the flight monitoring information acquired in the current data collection period with the time limit generated by the arbitrary positioning node information corresponding to the flight in the overrun event database.

[0023] S104: If the currently collected flight monitoring information meets the time generation strategy information restricted by any positioning node information, then t2 and t3 of the event to be monitored corresponding to the positioning node information are generated.

[0024] Specifically, the time generation strategy information in the prompt node information and positioning node information can be obtained according to the following steps: S110: Using the flight identity information (such as the flight number) included in the flight operation information, determine the real-time cabin voice information and QAR information corresponding to the current flight.

[0025] S120: Determine the operation description information identified from the real-time cabin voice information corresponding to the current flight as the aircraft's air-ground dialogue positioning information.

[0026] In actual usage scenarios, over-limit events of multiple flights will be monitored simultaneously. Therefore, it is necessary to use the flight number to determine the real-time cabin voice information and QAR information corresponding to the current flight from multiple cabin voice information and QAR information, so as to facilitate subsequent accurate matching.

[0027] S130: Use the own-air-ground dialogue positioning information to match the subsequently acquired real-time QAR information of the current flight.

[0028] During actual flight operations, pilots often communicate with ground-based air traffic control to request or confirm certain operations. For example, if a flight plan change (such as adjusting altitude or route) is required due to weather conditions, air traffic, or other reasons, the pilot must communicate with the relevant air traffic control unit and obtain approval. Once approved, the corresponding operation is performed, and a QAR (Qualified Air Traffic Response) message is generated in a short period of time.

[0029] Based on the above characteristics, this embodiment will use the operation description information identified by the real-time cabin voice information to determine in advance the timing of matching the aircraft's air-ground dialogue positioning information with the subsequent real-time QAR information of the current flight.

[0030] S140: If the operation information in the real-time QAR information is the same as the operation description information in the local air-ground dialogue positioning information, the generation time of the real-time QAR information is used as the positioning time, and the positioning time is used to generate t1, t2 and t3 corresponding to the event to be monitored.

[0031] In this embodiment, flight operation information (such as the flight number) is combined with real-time cabin voice recognition to locate the aircraft's air-ground conversation. This conversation is then combined with the QAR to generate a fix time. These fix times are key dependencies in the time generation strategy for generating t1, t2, or t3. Specifically, the corresponding time generation strategy information can be generated based on the logic between the fix time and t1, t2, and t3.

[0032] In addition, flight operation information and historical QAR engineering values ​​can be used to construct time-generated strategy information for each alert node or positioning node corresponding to each monitored event in the overrun event database. The generation of t1, t2, or t3 depends on the QAR information corresponding to a key node.

[0033] According to the judgment QAR information corresponding to the key nodes, t1, t2 and t3 corresponding to each monitored event can be generated in a timely and accurate manner in actual use. Figure 2 The QAR retrieval moment in .

[0034] In this embodiment, two key nodes are set to more clearly demonstrate the entire execution process of the monitored event. The prompt node (generated from the prompt node information) is typically a node before the monitored event and is used to inform the monitoring personnel which out-of-limit event the aircraft pilot will be performing in the future. The prompt node is used to generate t1. The positioning node (generated from the positioning node information) is typically the node at which the standard announcement operation should occur in the monitored event and is used to generate t2 and t3.

[0035] Taking the monitored event corresponding to the takeoff roll call at 80 knots as an example, another logic for establishing the overrun event database in this embodiment is described: Key node definition The airspeed reaching 60 knots is used as the prompt node corresponding to the event to be monitored to trigger a prompt. The corresponding prompt node information is the generation time of the QAR engineering value indicating that the airspeed reaches 60 knots, which is used as t1.

[0036] The airspeed reaching 80 knots is used as the positioning node corresponding to the event to be monitored to trigger the determination of the positioning start time and positioning end time. The corresponding positioning node information is the generation time of the QAR engineering value indicating that the airspeed reaches 80 knots. By combining different time offsets △T1 and △T2, t2 and t3 are generated.

[0037] Specifically, in a scenario that meets the requirements, the steps for generating t1, t2, and t3 are as follows: S101: If the aircraft QAR information acquired in the current data collection period is the same as the judgment QAR information corresponding to the key node in the prompt node information corresponding to any event to be monitored in the over-limit event database, the generation time of the aircraft QAR information is determined to be t1 corresponding to the event to be monitored.

[0038] S102: If the aircraft QAR information acquired during the current data collection period is identical to the judgment QAR information corresponding to the key node in the positioning node information corresponding to any event to be monitored in the overrun event database, then t2 and t3 corresponding to the event to be monitored are determined based on the generation time E of the aircraft QAR information. t2 and t3 satisfy the following conditions: t2 = E - △T1, t3 = E + △T2. △T1 and △T2 are the first and second time intervals, respectively. Because the length of voice messages varies in practice, and different drivers have different starting and ending rhythms, a pre-interval △T1 and a post-interval △T2 are set to maximize the time it takes to capture the voice messages in a comprehensive manner.

[0039] For example, for the monitoring event corresponding to the 80-knot roll call, the positioning start time is 4 seconds before the airspeed reaches 80 knots, that is, △T1 = 4 seconds, which serves as the starting time for the over-limit event judgment. The positioning end time is 5 seconds after the airspeed exceeds 80 knots, that is, △T2 = 5 seconds, which serves as the end node of the monitoring period.

[0040] In addition, when constructing the over-limit event database, the standard operation information to be executed for each monitored event in the over-limit event database can be constructed according to the SOP standard call content, such as the time range, call content, call role, and the number of times the call occurs.

[0041] In this step, by comparing the QAR data segments downloaded by the airborne system in real time with the over-limit event database, different over-limit determination periods corresponding to each monitored event can be dynamically activated in real time, and the corresponding over-limit monitoring tasks can be initiated. Specifically, in this embodiment, the over-limit determination period includes the determination prompt period [t1, t2] and the over-limit determination period (t2, t3]. Different tasks are performed in different periods, as explained below: S200: If there is a new complete and valid speech recognition result in the real-time cabin voice information collected in the current data collection cycle, traverse t1, t2 and t3 corresponding to each event to be monitored.

[0042] In this embodiment, only when a new speech recognition result is obtained will t1, t2 and t3 corresponding to each monitored event be traversed and the corresponding action execution strategy be started, thereby reducing the number of executions of the determination strategy and the amount of calculation.

[0043] S300: If the end time T of the last complete sentence in the speech recognition results of all valid voice channels obtained in the current data collection period and the t1 and t2 corresponding to any event to be monitored meet the condition T∈(t1, t2]), then the standard voice announcement information required by the SOP corresponding to the event to be monitored is generated.

[0044] Typically, an aircraft cabin is equipped with multiple audio channels, such as the pilot's headset, co-pilot's headset, and the cockpit audio receiver. Therefore, the spoken content appears in multiple different voice channels. To more comprehensively cover the recognition content of all voice channels and improve recognition accuracy, in this example, T is the end time of the last complete sentence in the speech recognition results of all valid voice channels in the current data collection cycle.

[0045] Shout prompts and preliminary judgment: If the current time exceeds the preset prompt start time or positioning start time, the upcoming shout content will be displayed on the interactive interface.

[0046] S400: If T and t2 and t3 corresponding to any event to be monitored satisfy the condition T∈(t2, t3]), then each complete speech recognition information obtained in the time window [t2, T] is compared with the standard voice call information required in the SOP corresponding to the event to be monitored.

[0047] The time window from the location start time (if not yet generated, the start time will be prompted instead) to the current time is used to verify the eligibility of the voice content, and the qualified content is displayed in real time on the interface. Because some monitoring events may contain multiple voice content, during this period, not only will the voice comparison test be performed, but the current qualified content will also be displayed in real time on the interface.

[0048] Specifically, in this step, the real-time voice information can be converted into text. Then, regular expressions are used to match and compare the standard voice announcement information with the two texts converted from the real-time voice information. The regular expressions for different announcement content matching rules can be flexibly set according to specific needs. For example, if both "Pre-taxi checklist" and "Pre-taxi inspection" are compliant, it can be written as "Pre-taxi inspection (list)?". To exclude "Pre-taxi checklist completed", it can be further optimized to "Pre-taxi inspection (list)?(?!completed)".

[0049] In order to make a more detailed and comprehensive judgment on the possibility of exceeding the limit in an over-limit event, the warning start time t4 can also be set to trigger the over-limit warning in advance. Usually t4∈(t2, t3], which can be configured according to specific needs.

[0050] Over-limit warning triggering: When an over-limit event has a warning start time configured and the current time has exceeded that time point, if the pass rate of the announcement content for the inspection item group falls below the preset threshold, the over-limit event warning mechanism will be immediately triggered. For example, the verification results of all announcement content within the time window will be fully displayed, allowing operators to quickly grasp the full situation.

[0051] Furthermore, in actual use scenarios, there are often similar events to be monitored. For example, during an aircraft's ascent, an announcement corresponding to each altitude value must be made, such as at 100 meters, 150 meters, and 200 meters. However, the intervals between these announcements actually shorten as the aircraft's ascent speed increases, resulting in the announcement content of a subsequent event to be monitored appearing within the interval [t2, t3] of the previous event to be monitored. Consequently, once the announcement content of the subsequent event to be monitored has appeared, the content of the previous event to be monitored must also exist. Therefore, in this embodiment, step S410 is provided to promptly update the original t3 corresponding to the previous event to be monitored to a more recent time when this occurs. Accordingly, the voice comparison time period for the previous event to be monitored can be promptly shortened, thereby reducing the amount of data processing required for voice recognition and comparison.

[0052] S410: During the voice information comparison process corresponding to any two adjacent similar monitored events, if the standard voice announcement information required by the SOP for the subsequent monitored event is identified within the time window [t2, t3] corresponding to the previous monitored event, t3 corresponding to the previous monitored event is updated to the time when the standard voice announcement information required by the SOP for the subsequent monitored event is identified. Similar monitored events are those of a type for which the similarity between the standard voice announcement information required by the SOP for the monitored events is greater than a preset similarity threshold.

[0053] S500: If T and t3 corresponding to any event to be monitored satisfy the condition T>t3, and in all voice comparison results of the event to be monitored within the time window [t2, t3], there is any voice recognition information that is the same as the standard voice call information required in the SOP corresponding to the event to be monitored, then the event to be monitored is determined to be qualified, that is, the current event to be monitored does not exceed the limit.

[0054] S510: If, in all voice comparison results of the event to be monitored within the time window [t2, t3], all voice recognition information is different from the standard voice announcement information required in the SOP corresponding to the event to be monitored, then the event to be monitored is determined to be an over-limit event of the flight.

[0055] Positioning end time: As the end node of the monitoring cycle, the judgment of over-limit events is stopped at this time point, and the final judgment result is output based on the monitoring results from the positioning start time to the positioning end time.

[0056] This embodiment is described using the monitored event corresponding to the takeoff roll announcement at 80 knots as an example: (1) Prompt trigger: Once the airspeed reaches 60 knots, the takeoff roll 80 knots announcement will be displayed on the operation interface.

[0057] (2) Dynamic judgment: When the end time of the last complete sentence in the speech recognition result exceeds 60 knots of airspeed, the following processing is performed: If the time node "4 seconds before airspeed 80 knots" has been generated, use this as the starting point for qualification judgment.

[0058] If the time node has not yet been generated, the content of the announcement can also be checked for eligibility with the airspeed of 60 knots as the starting point, and the qualified results that have been determined will be displayed in real time.

[0059] (3) Final judgment: When the end time of the last complete sentence in the speech recognition results of all valid channels exceeds "5 seconds after the airspeed reaches 80 knots", a comprehensive qualification judgment will be performed within the complete 9-second monitoring period (4 seconds before 80 knots to 5 seconds after 80 knots), and the final judgment result of the over-limit event will be generated.

[0060] This invention enables real-time reception and processing of multi-source aircraft (i.e., aircraft) data. The out-of-limit event identification method, based on this multi-source real-time data, makes identification of out-of-limit events more accurate and timely. Compared to traditional post-analysis methods, this method can detect out-of-limit events during flight, significantly improving the timeliness of flight risk warnings, buying valuable time for emergency measures, and effectively reducing the likelihood of flight accidents.

[0061] In addition, as another possible embodiment of the present invention, the following method is also provided: S501: If the event to be monitored corresponds to only one set of prompt node information and positioning node information, the state of the event to be monitored is configured as a monitored event.

[0062] S502: If the event to be monitored corresponds to multiple groups of prompt node information and positioning node information, the positioning end time of the group that has completed monitoring is updated as the historical monitoring end time.

[0063] S503: If the newly acquired prompt start time of the event to be monitored is greater than the historical monitoring deadline, the monitoring task corresponding to the newly acquired prompt start time is started.

[0064] S504: If the newly acquired prompt start time of the event to be monitored is less than the historical monitoring deadline, the monitoring task corresponding to the newly acquired prompt start time is not started.

[0065] For events to be monitored that only have one set of prompt node information and positioning node information, after the judgment is completed, the monitoring time period will be automatically recorded in the historical archive of the corresponding call item, and the status of the event to be monitored will be changed to a monitored event. In the subsequent per-second cycle monitoring, the time period for which monitoring has been completed will be automatically skipped to avoid repeated detection and improve monitoring efficiency.

[0066] For events to be monitored that have multiple sets of prompt node information and positioning node information, after each corresponding detection content is completed, the positioning end time of the currently monitored group will be automatically updated to the historical monitoring deadline. In the subsequent per-second cyclic monitoring, the newly obtained prompt start time of the event to be monitored can be used to avoid repeated detection and improve monitoring efficiency based on whether it is greater than the historical monitoring deadline.

[0067] As another possible embodiment of the present invention, there are multiple voice channels, such as Figure 3 As shown, the effective voice channel is obtained according to the following steps: S600: When the complete voice information of any voice channel is recognized for the first time, which is the standard voice call information required in the SOP corresponding to the event to be monitored, the generation period of the complete voice information of the voice channel is used as the determination period.

[0068] S700: The voice information generated in other voice channels within the determination period is used as the determined voice information.

[0069] S800: Obtaining the similarity between each determined voice information and the complete voice information of the first recognized voice channel.

[0070] Specifically, the similarity between the voice information and the complete voice information of the first recognized voice channel is determined and obtained according to the following steps: S801: performing frame processing on the determined voice information and the complete voice information of the first recognized voice channel respectively.

[0071] Frame processing: Continuous speech signals are divided into small time segments (typically 20-30 milliseconds). Because speech signals vary over time, they are inherently non-stationary, meaning their statistical properties (such as frequency content) vary over time. However, over short timescales (e.g., 20-40 milliseconds), speech signals can be assumed to be "quasi-stationary." By dividing the speech signal into short time segments (frames), analysis methods suitable for stationary signals can be applied to each small segment.

[0072] Specifically, before S801, other audio preprocessing operations can be set, such as pre-emphasis: using a high-pass filter to enhance high-frequency components to compensate for high-frequency components lost during pronunciation. After S801, windowing can also be set: applying a Hamming window or other window function to each frame to reduce spectral leakage.

[0073] S802: According to the frequency spectrum information corresponding to each frame, respectively obtain the Mel-frequency cepstral coefficients corresponding to the determined voice information and the complete voice information of the first recognized voice channel.

[0074] S803: Based on the Mel-frequency cepstral coefficients, a dynamic time warping method is used to obtain the similarity between the determined voice information and the complete voice information of the first recognized voice channel.

[0075] Extract the MFCC coefficients (also known as Mel-frequency cepstral coefficients) of speech A and speech B to obtain two matrices X∈R T1×d and Y∈R T2×d , where T1 and T2 are the number of frames, and d is the MFCC dimension.

[0076] Construct a distance matrix D, where , D ij represents the Euclidean distance between the i-th frame of speech A and the j-th frame of speech B, X i ∈X,Y j ∈Y.

[0077] Dynamic Time Warping (DTW) is used to find the shortest path, and the sum of the distances on the accumulated path is used as the similarity score (the smaller the distance, the more similar the paths).

[0078] The selection of the shortest path must meet the following conditions: The path must start from the upper left corner of the matrix and reach the lower right corner.

[0079] The path can only be moved downward, rightward, or diagonally (i.e., it cannot be moved backward).

[0080] There is a constraint that limits the slope of the path, ensuring temporal continuity and locality.

[0081] S900: When the similarity between any determined voice information and the complete voice information of the first recognized voice channel is greater than a preset similarity threshold, the voice channel to which the determined voice information belongs is determined as a valid voice channel.

[0082] Because multiple voice channels in the cabin may receive the same voice content due to different transmission distances and different performance of encoding and decoding equipment, the same voice content may have a certain delay in different voice channels. Directly comparing their features may lead to inaccurate similarity calculations. In order to accurately calculate the similarity of two audios in this case, the DTW technology is used in this embodiment to calculate the similarity between voice audios. Since DTW allows nonlinear stretching or compression on the time axis to find the best matching path, it can be used to measure the similarity of two time series under nonlinear alignment on the time axis. It can then be applied to the present embodiment to more accurately compare the similarity between voice signals of different lengths.

[0083] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0084] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, mobile terminal, or network device) to execute the methods according to the embodiments of the present disclosure.

[0085] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0086] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "systems."

[0087] The electronic device according to this embodiment of the present invention is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0088] Electronic devices are presented in the form of general-purpose computing devices. Figure 4 As shown, the components of the electronic device may include but are not limited to: the at least one processor mentioned above, the at least one memory mentioned above, and a bus connecting different system components (including the memory and the processor).

[0089] The storage stores program codes, which can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0090] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read only memory (ROM).

[0091] The storage may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0092] The bus may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures.

[0093] The electronic device may also communicate with one or more external devices (e.g., a keyboard, pointing device, Bluetooth device, etc.) via a communication interface, one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., a router, modem, etc.). This communication may occur via an input / output (I / O) interface. Furthermore, the electronic device may communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0094] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, terminal device, or network device) to execute the methods according to the embodiments of the present disclosure.

[0095] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0096] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0097] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0098] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0099] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0100] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0101] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0102] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for monitoring over-limit events based on multi-source real-time data, characterized in that: The method includes the following steps: Match the flight monitoring information obtained in the current data collection period with the prompt node information and positioning node information corresponding to each to-be-monitored event in the over-limit event database to generate the prompt start time t1, positioning start time t2, and positioning end time t3 corresponding to each to-be-monitored event; t1 < t2 < t3; the to-be-monitored event is a standard operation task including a voice call task in the SOP corresponding to the current flight; the prompt node information and positioning node information are respectively time generation strategy information formed by cabin sound information and / or QAR information generated when the preset operations corresponding to the to-be-monitored event are executed; the flight monitoring information includes one or any combination of flight operation information, real-time cabin sound information, and real-time QAR information; If there is a new complete and valid speech recognition result in the real-time cabin sound information collected in the current data collection period, traverse t1, t2, and t3 corresponding to each to-be-monitored event; If the end time T of the last complete sentence in the speech recognition results of all valid voice channels obtained in the current data collection period satisfies the condition T ∈ (t1, t2] for any to-be-monitored event, generate the standard voice call information required in the SOP corresponding to the to-be-monitored event; If T satisfies the condition T ∈ (t2, t3] for any to-be-monitored event corresponding to t2 and t3, compare each complete speech recognition information obtained in the time window [t2, T] with the standard voice call information required in the SOP corresponding to the to-be-monitored event; If T satisfies the condition T > t3 for any to-be-monitored event, and if in all speech comparison results of the to-be-monitored event in the time window [t2, t3], all speech recognition information is different from the standard voice call information required in the SOP corresponding to the to-be-monitored event, then determine that the to-be-monitored event is an over-limit event of the flight.

2. The method according to claim 1, characterized in that After T satisfies the condition T > t3 for any to-be-monitored event, the method further includes: If there is any speech recognition information in all speech comparison results of the to-be-monitored event in the time window [t2, t3] that is the same as the standard voice call information required in the SOP corresponding to the to-be-monitored event, then determine that the to-be-monitored event is qualified.

3. The method according to claim 2, characterized in that After determining whether the to-be-monitored event is over-limit, the method further includes: If the to-be-monitored event only corresponds to a set of prompt node information and positioning node information, configure the status of the to-be-monitored event as a monitored event; If the to-be-monitored event corresponds to multiple sets of prompt node information and positioning node information, update the positioning end time of the currently completed monitored set to the historical monitoring cut-off time; If the prompt start time of the newly obtained to-be-monitored event is greater than the historical monitoring cut-off time, start the monitoring task corresponding to the newly obtained prompt start time; If the prompt start time of the newly obtained to-be-monitored event is less than the historical monitoring cut-off time, do not start the monitoring task corresponding to the newly obtained prompt start time.

4. The method according to claim 1, wherein Match the flight monitoring information acquired during the current data collection period with the prompt node information and positioning node information corresponding to each event to be monitored in the overrun event database to generate the prompt start time t1, positioning start time t2, and positioning end time t3 corresponding to each event to be monitored; including: Matching flight monitoring information acquired during the current data collection cycle with the time-limited generation strategy information of any corresponding prompt node information of the flight in the overrun event database; If the currently collected flight monitoring information meets the time generation strategy information limited by any prompt node information, then generate t1 of the event to be monitored corresponding to the prompt node information; Matching flight monitoring information acquired during the current data collection cycle with the time-limited generation strategy information of any positioning node corresponding to the flight in the overrun event database; If the currently collected flight monitoring information satisfies the time generation strategy information restricted by any positioning node information, then t2 and t3 corresponding to the event to be monitored of the positioning node information are generated.

5. The method according to claim 4, characterized in that The time generation strategy information includes: Using the flight identity information included in the flight operation information, determine the real-time cabin voice information and QAR information corresponding to the current flight; Determine the operation description information identified from the real-time cabin voice information corresponding to the current flight as the aircraft's air-ground dialogue positioning information; Use the own-aircraft ground dialogue positioning information to match the subsequently acquired real-time QAR information of the current flight; If the operation information in the real-time QAR information is the same as the operation description information in the local air-ground dialogue positioning information, the generation time of the real-time QAR information is used as the positioning time, and the positioning time is used to generate t1, t2 and t3 corresponding to the event to be monitored.

6. The method according to claim 1, characterized in that Also includes: In the voice information comparison processing corresponding to any two adjacent similar events to be monitored, if the standard voice call information required in the SOP corresponding to the subsequent event to be monitored is identified within the time window [t2, t3] corresponding to the previous event to be monitored, then t3 corresponding to the previous event to be monitored is updated to the time when the standard voice call information required in the SOP corresponding to the subsequent event to be monitored is identified; similar events to be monitored are a type of events to be monitored whose similarity with the standard voice call information required in the SOP corresponding to the event to be monitored is greater than a preset similarity threshold.

7. The method according to claim 1, characterized in that There are multiple voice channels. To obtain a valid voice channel, follow the steps below: When the complete voice information of any voice channel is recognized for the first time, which is the standard voice call information required in the SOP corresponding to the event to be monitored, the generation period of the complete voice information of the voice channel is used as the determination period; using the voice information generated in the other voice channels within the determination period as the determined voice information; Obtaining the similarity between each determined voice information and the complete voice information of the first recognized voice channel; When the similarity between any determined voice information and the complete voice information of the first recognized voice channel is greater than a preset similarity threshold, the voice channel to which the determined voice information belongs is determined as a valid voice channel.

8. The method according to claim 7, characterized in that Determine the similarity between the voice information and the complete voice information of the first recognized voice channel, and obtain it according to the following steps: The judgment voice information and the complete voice information of the first recognized voice channel are framed separately; According to the spectrum information corresponding to each frame, the Mel-frequency cepstral coefficients corresponding to the judgment voice information and the complete voice information of the first recognized voice channel are obtained respectively; Based on the Mel-frequency cepstral coefficients, the dynamic time warping method is used to obtain the similarity between the determined speech information and the complete speech information of the first recognized speech channel.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for monitoring over-limit events based on multi-source real-time data as described in any one of claims 1 to 8 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for monitoring over-limit events based on multi-source real-time data as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Traffic over-limit alarm method and system

    CN115035747A

  • Aircraft overrun identification method based on multi-source sensing information fusion

    CN117874628A

  • Abnormal flight overrun event identification method and system

    CN120199113A

  • Data alignment method and device

    CN120388432A

  • Premonition capture and effectiveness evaluation method and system for monitoring aviation unsafe events in real time

    CN120496369A