An intelligent flight training system
The intelligent flight training system simulates incorrect commands from the pilot by adjusting the virtual voice file and volume, solving the problem of the lack of real-world scenarios in single-pilot training and achieving a more realistic and quantifiable training effect.
Patent Information
- Application Number
- CN202510704449.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing flight training methods are insufficient to realistically simulate erroneous commands issued by the pilot and assess the co-pilot's error correction capabilities in single-pilot training, especially in virtual voice command scenarios where effective simulation methods are lacking.
The intelligent flight training system utilizes a voice playback device and processor to adjust the content and volume of the virtual voice file based on the co-pilot's command time within a preset sliding time window. This simulates erroneous commands and reaction times in real-world scenarios. Combined with reverse command processing methods, the training effect is quantitatively evaluated.
This technology enables more realistic simulation of incorrect commands from the driver during single-person co-driver training, improving the consistency and quantitative evaluation capabilities of the training and enhancing its effectiveness.
Smart Images

Figure CN120319094B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent flight training technology, and in particular to an intelligent flight training system. Background Technology
[0002] In pilot training, the process is divided into pilot training and co-pilot training. Typically, the co-pilot executes the instructions issued by the pilot. However, current flight training methods are usually geared towards multi-crew training, requiring the pilot and co-pilot to work together, or an instructor to assist the student in completing the training. However, in some situations, single-pilot training is necessary. To address this, virtual methods are often used in conjunction with single-pilot training. That is, the pilot issues virtual voice commands, and the co-pilot executes the corresponding actions. In real-world scenarios, the pilot may issue incorrect commands for various reasons. The co-pilot then needs to determine if the pilot's commands are problematic and assess the co-pilot's ability to correct errors. How to realistically simulate these scenarios during single-pilot training has become a pressing technical problem. Summary of the Invention
[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0004] According to the intelligent flight training system provided in this application, the system includes: a voice playback device, a storage medium, and a processor; wherein, the voice playback device is used to play virtual voice corresponding to the pilot; the storage medium stores at least one instruction or at least one program segment, which is loaded and executed by the processor to achieve the following steps:
[0005] H100, if the target user is in the passenger seat, then obtain the time taken for each virtual voice command executed by the target user within the preset sliding time window TH, to obtain a time list T = (T1, T2, ..., T...). x ,…,T y ), x = 1, 2, ..., y; where, T x The time taken to execute the instruction corresponding to the xth virtual voice within TH for the target user, where y is the number of virtual voice instructions executed by the target user within TH.
[0006] H200, based on T, determine the average time T' = (1 / y) × ∑ for the target user to execute the virtual voice command within TH. y x=1 T x .
[0007] H300, if T' < T1 or T' > T2, then obtain the next virtual voice file YA to be played; where T1 is the first preset time consumption threshold and T2 is the second preset time consumption threshold; T1 < T2.
[0008] H400, if there is a preset reverse instruction for the instruction corresponding to YA, then the first preset virtual voice processing method is used to process YA to obtain the reverse virtual voice file YB corresponding to YA.
[0009] H500 plays YB through the voice playback device and obtains the actual instruction YB' corresponding to the target user's actual operation on YB.
[0010] H600 determines the target user's completion rate for YA based on the instructions corresponding to YB' and YA.
[0011] The present invention has at least the following beneficial effects:
[0012] The intelligent flight training system of the present invention, if the target user is in the co-pilot position, determines the average time taken for the target user to execute the virtual voice commands within the preset sliding time window TH based on the time taken for each virtual voice command executed by the target user within TH. If the average time taken is less than a first preset time or greater than a second preset time, the next virtual voice file to be played is obtained. If the next virtual voice file to be played contains a preset reverse command, the next virtual voice file to be played is processed using a first preset virtual voice processing method to obtain the corresponding reverse virtual voice file. The reverse virtual voice file is played through the voice playback device. Since the reverse virtual voice file is generated based on the duration of the target user's command execution and does not have a certain regularity, it makes the simulation of the target user's actions against the co-pilot more consistent with the real flight scenario.
[0013] Furthermore, after the voice playback device plays the reverse virtual voice command, it also obtains the actual command corresponding to the target user's actual operation of the reverse virtual voice command, thereby determining the target user's completion rate for the next voice file to be played, and realizing a quantitative evaluation of the target user's training. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 A flowchart illustrating the steps executed by the processor in the intelligent flight training system provided in this embodiment of the invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] It should be noted that, based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Furthermore, this device and / or practice the method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.
[0018] Example 1:
[0019] The following will refer to Figure 1 The flowchart shown illustrates the steps executed by the processor in an intelligent flight training system, introducing such a system.
[0020] The intelligent flight training system includes: a voice playback device, a storage medium, and a processor; wherein, the voice playback device is used to play the virtual voice corresponding to the pilot.
[0021] In this embodiment, the intelligent flight training system can conduct flight simulation training for a single person. During single-person training, the target user can be either the pilot or the co-pilot. The intelligent flight training system is equipped with a voice playback device, which can simulate the pilot's voice. When the target user is the co-pilot, they can execute corresponding instructions based on the voice played by the voice playback device.
[0022] The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to perform the following steps:
[0023] H100, if the target user is in the passenger seat, then obtain the time taken for each virtual voice command executed by the target user within the preset sliding time window TH, to obtain a time list T = (T1, T2, ..., T...). x ,…,T y ), x = 1, 2, ..., y; where, T x The time taken to execute the instruction corresponding to the xth virtual voice within TH for the target user, where y is the number of virtual voice instructions executed by the target user within TH.
[0024] In this embodiment, if the target user is in the passenger seat, it means that the target user is conducting simulation training for the passenger. Within the sliding time window TH, the target user executes several instructions corresponding to virtual voices played by the voice playback device. When executing each instruction corresponding to a virtual voice, the execution time of each instruction corresponding to a virtual voice can be recorded, and thus T can be obtained.
[0025] H200, based on T, determine the average time T' = (1 / y) × ∑ for the target user to execute the virtual voice command within TH. y x=1 T x .
[0026] In this embodiment, after obtaining the time taken for the target user to execute each virtual voice instruction within TH, T' can be obtained; T' can characterize the target user's current reaction ability; T' can be set according to actual needs or obtained based on experience, for example: T' = 1min.
[0027] H300, if T' < T1 or T' > T2, then obtain the next virtual voice file YA to be played; where T1 is the first preset time consumption threshold and T2 is the second preset time consumption threshold; T1 < T2.
[0028] In this embodiment, it should be noted that, under normal circumstances, the time taken for the target user to execute each virtual voice command will not vary significantly. If T' < T1, it indicates that the time taken for the target user to execute the virtual voice command is too short, and the target user may have anticipated the voice played by the virtual voice playback device based on experience and executed the corresponding command in advance. If T' > T2, it indicates that the time taken for the target user to execute the virtual voice command is too long, and the target user may be inattentive or fatigued, resulting in a slow reaction. In this case, the next virtual voice file YA to be played is obtained. T1 and T2 can be obtained based on experience or through extensive experimentation, and will not be elaborated here.
[0029] H400, if there is a preset reverse instruction for the instruction corresponding to YA, then the first preset virtual voice processing method is used to process YA to obtain the reverse virtual voice file YB corresponding to YA.
[0030] In this embodiment, YA can be converted into text, and then semantic recognition can be performed to obtain the corresponding instruction ZA; a reverse instruction mapping table is preset, which contains several standard instructions and the reverse instruction corresponding to each standard instruction; it can be determined whether the instruction corresponding to YA has a preset reverse instruction by traversing the reverse instruction mapping table.
[0031] Furthermore, step H400 may include the following steps:
[0032] H410, retrieve the instruction ZA corresponding to YA.
[0033] H420, retrieve the inverse instruction ZA' corresponding to ZA.
[0034] H430, extract keywords from ZA and ZA' to obtain the keyword list GA = (GA1, GA2, ..., GA') corresponding to ZA. u , ...,GA v ) and the keyword list corresponding to ZA' GA' = (GA'1, GA'2, ..., GA') u , ...,GA' v ),u=1,2,…,v; among them, GA u For the u-th keyword corresponding to ZA, GA' u Let v be the u-th keyword corresponding to ZA', and v be the number of keywords corresponding to ZA and ZA'.
[0035] In this embodiment, the keywords can be words other than auxiliary words and modifiers; those skilled in the art can use existing keyword extraction methods to extract keywords from ZA and ZA' according to actual needs, which will not be elaborated here.
[0036] H440, iterate through GA and GA', if GA u with GA' u If they are different, then GA' u The inverse keyword corresponding to ZA' has been identified.
[0037] In this embodiment, the standard command and its corresponding reverse command differ in keywords. For example, the reverse command for turning 30° to the left is turning 30° to the right; if GA u with GA' u If they are different, it means GA' u This is the inverse keyword corresponding to ZA'.
[0038] H450, adjust the volume of the audio file segment corresponding to the reverse keyword to the first volume YL1 to obtain YB; where YL1=λ×YL; YL is the volume preset by the target user; λ is the preset reverse keyword volume weight, λ<1.
[0039] In this embodiment, the reverse instruction corresponds to virtual voice, and the reverse keyword in the reverse instruction corresponds to a voice segment in the virtual voice. The volume of the reverse voice segment can be turned down to obtain YB. The preset volume for the target user is the volume at which the target user can hear the virtual voice clearly. λ can be set according to actual needs or determined through a large number of experiments. For example, the value range of λ is 0.2 to 0.5.
[0040] Furthermore, after step H400 and before step H500, the at least one instruction or the at least one program segment is loaded and executed by the processor, and the following steps are also implemented:
[0041] H460, if there is no preset reverse instruction for the instruction corresponding to YA, then the second preset virtual voice processing method is used to process YA to obtain the reverse virtual voice file YB corresponding to YA.
[0042] In this embodiment, the virtual voice may also correspond to non-standard commands, so that the command corresponding to YA does not have a preset reverse command. In this case, the second preset virtual voice processing method is used to process YA. Specifically, H460 may include the following steps:
[0043] H461, retrieve the volume of each preset user within a preset historical time period to obtain a historical preset volume list HE = (HE1, HE2, ..., HE...). a HE b ), a = 1, 2, ..., b; where, HE a b represents the a-th preset volume value for the target user within a preset historical time period, and b represents the number of preset volume values for the target user within the preset historical time period.
[0044] In this embodiment, the target user will have a corresponding historical preset volume in each historical training process, and the preset volume of the target user within the preset historical time period can be obtained to obtain HE.
[0045] H462 uses a preset clustering algorithm to cluster the volume within HE to obtain several clusters.
[0046] H463 determines the average volume corresponding to the cluster with the most volume as the target volume ML corresponding to YA.
[0047] In this embodiment, the volume of the cluster with the most volume represents the volume that the target user frequently uses, which is consistent with the target user's volume setting habits. Therefore, the average volume corresponding to the cluster with the most volume is determined as the target volume corresponding to YA.
[0048] H464, adjust the volume of YA to the second volume YL2 to obtain YB; where YL2 = η × ML; η is the preset volume adjustment weight; η < 1.
[0049] In this embodiment, the value of η can be between 0.3 and 0.4; for example, η = 0.35; the volume of YA is turned down to simulate a real scenario where the passenger is not in good condition and the driver speaks softly, thereby training the target user's ability to handle situations as a passenger.
[0050] H500 plays YB through the voice playback device and obtains the actual instruction YB' corresponding to the target user's actual operation on YB.
[0051] In this embodiment, the virtual voice played by the voice playback device may be incorrect or correct voice at a low volume, thereby simulating a real scene without any regularity, thus improving the training effect.
[0052] H600 determines the target user's completion rate for YA based on the instructions corresponding to YB' and YA.
[0053] Furthermore, step H600 may include the following steps:
[0054] H610, if the instruction corresponding to YB' is the same as that corresponding to YA, then obtain the duration TG from the end of YB playback to the target user completing YB'.
[0055] In this embodiment, if the instructions corresponding to YB' and YA are the same, it means that the target user has accurately identified the instructions corresponding to the virtual voice; the duration from the end of YB playback to the target user completing YB' can be obtained.
[0056] Furthermore, if the instructions corresponding to YB' and YA are different, it means that the target user failed to accurately identify the instructions corresponding to the virtual voice, and the execution of this instruction is unqualified.
[0057] H620, if TG∈[TG1,TG2], then determine the completion degree DQ of the target user for YA = (TG1 / 2 + TG2 / 2) / TG; where TG1 is the preset minimum time to complete the instruction and TG2 is the preset maximum time to complete the instruction.
[0058] In this embodiment, TG1 and TG2 can be set according to actual needs or obtained through a large number of experiments; the larger the DQ, the higher the target user's completion rate, and the smaller the DQ, the lower the target user's completion rate.
[0059] H630, if TG < TG1 or TG > TG2, then the target user's completion rate DQ for YA is determined to be 0.
[0060] In this embodiment, the duration for the target user to correctly execute the virtual voice command should be within a preset duration range; both excessively long and excessively short durations indicate abnormal execution by the target user.
[0061] Furthermore, after step H600, the at least one instruction or the at least one program segment is loaded and executed by the processor, and the following steps are also implemented:
[0062] H700, obtain the completion rate of the target user's command for each virtual voice, to obtain a completion rate list FE = (FE1, FE2, ..., FE2). c , ..., FE d ), c = 1, 2, ..., d; where, FE c Let d represent the completion rate of the target user's command for the c-th virtual voice, and d represent the number of virtual voices.
[0063] H710, based on FE, determine the average completion rate FE' = (1 / d)∑ for the target user. d c=1 FE c .
[0064] H720, if FE'≥FR, then the target user is deemed to have passed the training; otherwise, the target user is deemed to have failed the training.
[0065] In this embodiment, the above steps can be used to determine the average completion rate of the target user's training, thereby specifically quantifying whether the target user's simulated training is qualified.
[0066] In this embodiment of the intelligent flight training system, if the target user is in the co-pilot position, the system determines the average time taken for the target user to execute the virtual voice commands within the preset sliding time window TH, based on the time taken for each virtual voice command executed by the target user within TH. If the average time taken is less than a first preset time or greater than a second preset time, the system obtains the next virtual voice file to be played. If the next virtual voice file to be played contains a preset reverse command, the system processes the next virtual voice file using a first preset virtual voice processing method to obtain the corresponding reverse virtual voice file. The reverse virtual voice file is then played through the voice playback device. Since the reverse virtual voice file is generated based on the duration of the target user's command execution and does not have a certain regularity, it makes the simulation of the target user's actions against the co-pilot more consistent with real flight scenarios.
[0067] Furthermore, after the voice playback device plays the reverse virtual voice command, it also obtains the actual command corresponding to the target user's actual operation of the reverse virtual voice command, thereby determining the target user's completion rate for the next voice file to be played, and realizing a quantitative evaluation of the target user's training.
[0068] Example 2:
[0069] Based on Example 1, the target user may also conduct simulation training for the driver's seat. When the target user conducts simulation training for the driver's seat, the following method is used:
[0070] S100, if the target user is in the pilot's position, then acquire the target voice sent by the target user; wherein, the target user is any user conducting flight training.
[0071] In this embodiment, it is understood that when conducting simulated flight training for a target user, the target user can train as either the pilot or the co-pilot. If the target user is in the pilot's position, the target user is conducting simulated training as the pilot. Under normal circumstances, the pilot issues voice commands to the co-pilot, and the co-pilot executes the commands corresponding to the pilot's voice commands. When the target user is in the pilot's position and issues a voice command, the target voice command issued by the target user can be acquired. Whether the target user is in the pilot's position can be determined by image recognition or human detection sensors, which will not be elaborated here.
[0072] S200, the semantic information of the target speech is matched with each preset standard instruction to obtain a standard matching degree list corresponding to the target speech; wherein, the standard matching degree list includes the standard matching degree between the semantic information of the target speech and each preset standard instruction.
[0073] Specifically, the semantic information of the target speech is matched with each preset standard instruction to obtain a standard matching degree list A = (A1, A2, ..., A...). i A n ), i=1, 2,..., n; among them, A i Let n be the matching degree between the target speech and the i-th standard instruction, and n be the number of preset standard instructions.
[0074] In this embodiment, several standard instructions are preset, such as turning left 30°, climbing up 1000m, etc. The target speech can be converted into text information first, and then the corresponding speech information is obtained. The standard instructions also have corresponding speech information. It should be noted that those skilled in the art can use existing semantic similarity determination methods according to actual needs to match the semantic information of the target speech with each preset standard instruction. This will not be elaborated here.
[0075] S300, the standard matching degree in the standard matching degree list that is greater than or equal to the first preset matching degree threshold is determined as the target matching degree, so as to obtain the target matching degree list; wherein, the target matching degree list includes several target matching degrees.
[0076] Specifically, iterate through A, if A i If ≥α1, then A i Determine the target match degree; to obtain the target match list B = (B1, B2, ..., B...). j B m), j = 1, 2, ..., m; where, B j The j-th target matching degree is determined, m is the number of target matching degrees determined; α1 is the preset first matching degree threshold.
[0077] In this embodiment, if A i ≥α1 indicates that the semantic information of the i-th standard instruction is quite similar to the speech information corresponding to the target speech, and the i-th standard instruction is likely to be the standard instruction corresponding to the target speech.
[0078] Furthermore, the value range of α1 is [0.7, 0.8]; for example: α1 = 0.8; α1 can be obtained based on experience or through a large number of experiments.
[0079] S400, if the number of target matching degrees in the target matching degree list is greater than 1, then randomly determine the standard instruction corresponding to one of the target matching degrees in the target matching degree list as the target instruction corresponding to the target speech.
[0080] Specifically, if m > 1, then the standard instruction corresponding to a target matching degree in B is randomly determined as the target instruction corresponding to the target speech.
[0081] Understandably, normally a target speech corresponds to only one standard command. However, in this case, at least one standard command is identified, including both the actual command corresponding to the target speech and erroneous commands. Since the standard command corresponding to a target match degree in the target match degree list is randomly selected as the target command corresponding to the target speech, the identified target command may be incorrect. This simulates a real scenario where the co-pilot executes an incorrect command. Erroneous commands are generated randomly and without regularity, making it impossible for the target personnel to predict them in advance, thus improving the training effect.
[0082] S500, if the number of target matching degrees in the target matching degree list is equal to 1, and the only target matching degree in the target matching degree list is greater than or equal to the second preset matching degree threshold, then the standard instruction corresponding to the only target matching degree in the target matching degree list is determined as the target instruction corresponding to the target speech; the first preset matching degree threshold is less than the second preset matching degree threshold.
[0083] Specifically, if m = 1 and B1 ≥ α2, then the standard instruction ZL corresponding to B1 is determined as the target instruction corresponding to the target speech; where α2 is the preset second matching degree threshold; α2 > α1.
[0084] In this embodiment, if m = 1, B1 ≥ α2, and α2 > α1, it means that the only standard instruction determined is the same as the actual instruction corresponding to the target speech; this situation indicates that the actual instruction corresponding to the target speech has been determined, which is a relatively normal situation; this situation can be achieved when the target user speaks clearly and in a standard way.
[0085] Furthermore, the value range of α2 is [0.9, 0.95]; for example: α2 = 0.92; α2 can be obtained based on experience or through a large number of experiments.
[0086] S600, if the number of target matching degrees in the target matching degree list is equal to 1, and the only target matching degree in the target matching degree list is less than the second preset matching degree threshold, then the target speech is converted into the corresponding target text.
[0087] Specifically, if m = 1 and B1 < α2, then the target speech is converted into the corresponding target text.
[0088] In this embodiment, if m = 1 and B1 < α2, it indicates that the determined standard instruction may be the actual instruction corresponding to the target speech. Since B1 < α2, the similarity between the two is not high enough to be directly determined. Therefore, it is necessary to further determine whether the determined standard instruction is the actual instruction corresponding to the target speech. It should be noted that those skilled in the art can use existing speech-to-text methods to convert the target speech into the corresponding target text according to actual needs, which will not be elaborated here.
[0089] S700: Determine the target command corresponding to the target voice based on the target text, and execute the target command through the preset co-pilot control module.
[0090] Furthermore, step S700 may include the following steps:
[0091] S710, perform initial keyword extraction on the target text to obtain an initial keyword list corresponding to the target text; wherein, the initial keyword list includes several initial keywords corresponding to the target text.
[0092] Specifically, initial keywords are extracted from the target text to obtain an initial keyword list C = (C1, C2, ..., C...). p C q ), p = 1, 2, ..., q; where C p Let q be the p-th initial keyword obtained from the initial keyword extraction of the target file, and q be the number of initial keywords obtained from the initial keyword extraction of the target file.
[0093] In this embodiment, keywords can be words other than auxiliary words and modifiers; those skilled in the art can use existing keyword extraction methods to extract initial keywords from the target text according to actual needs, which will not be elaborated here.
[0094] S720, obtain the confidence level of each initial keyword in the initial keyword list to obtain the initial keyword confidence level list; wherein, the initial keyword confidence level list includes the confidence level of each initial keyword in the initial keyword list.
[0095] Specifically, obtain the confidence score corresponding to each initial keyword in C to obtain the initial keyword confidence score list TC = (TC1, TC2, ..., TC3) for C. P , ..., TC q ); where TC P C p The corresponding confidence level.
[0096] In this embodiment, the target text is obtained from the target speech. When generating the target text from the target speech, it is also generated word by word. Therefore, each word corresponds to a confidence level.
[0097] S730, iterate through the initial keyword confidence list, and determine the initial keywords in the initial keyword confidence list that are less than the preset keyword confidence threshold as intermediate keywords to obtain the intermediate keyword list.
[0098] Specifically, iterate through TC, if TC P <β, then C P The keywords are determined to be intermediate keywords, resulting in a list of intermediate keywords D = (D1, D2, ..., D...). r D s ), r = 1, 2, ..., s; where, D r The r-th intermediate keyword is determined, s is the number of determined intermediate keywords, and β is the preset keyword confidence threshold.
[0099] In this embodiment, keywords with low confidence may be incorrect keywords obtained from the target speech translation; for example, when the target user says "climbing 1000m upwards", the pronunciation of "upwards" is relatively unclear, and the confidence of the keyword "upwards" is relatively low.
[0100] Furthermore, the value of β can range from 0.5 to 0.7; for example, β = 0.6; β can be obtained empirically or through a large number of experiments.
[0101] S740, traverse the intermediate keyword list. If there is an intermediate keyword in the intermediate keyword list that belongs to the preset instruction keyword library, then determine the target instruction corresponding to the target speech through the preset reverse instruction mapping table. The reverse instruction mapping table includes several standard instructions and the reverse instruction corresponding to each standard instruction.
[0102] Specifically, iterate through D. If there is an intermediate keyword in D that belongs to the preset instruction keyword library, then determine the target instruction corresponding to the target speech through the preset reverse instruction mapping table QT.
[0103] In this embodiment, a pre-set instruction keyword library is provided, which contains several keywords corresponding to standard instructions. If there is an intermediate keyword in D that belongs to the pre-set instruction keyword library, it means that there is a keyword corresponding to the standard instruction among the determined intermediate keywords. That is, when the target user emits the target voice, the pronunciation of the key words is unclear, and the co-driver may not be able to hear the actual target voice clearly. At this time, the target instruction corresponding to the target voice is determined through the pre-set reverse instruction mapping table QT.
[0104] Furthermore, determining the target instruction corresponding to the target speech through a preset reverse instruction mapping table may include the following steps:
[0105] S741, traverse the reverse instruction mapping table, and determine the reverse instruction of the standard instruction mapping that is the same as the standard instruction corresponding to the only target matching degree in the target matching degree list as the target instruction corresponding to the target speech.
[0106] In this embodiment, the preset reverse instruction mapping table contains the reverse instruction corresponding to each standard instruction. For example, the reverse instruction corresponding to the standard instruction to climb 1000m upward is to descend 1000m downward. In step S200, a standard instruction is matched, and the reverse instruction corresponding to the standard instruction can be obtained through the preset reverse instruction mapping table.
[0107] Furthermore, determining the target instruction corresponding to the target speech based on the target text may further include the following steps:
[0108] S750, if there is no intermediate keyword in the intermediate keyword list that belongs to the preset instruction keyword library, then the standard instruction that is the same as the standard instruction corresponding to the only target matching degree in the target matching degree list in the reverse instruction mapping table is determined as the target instruction corresponding to the target speech.
[0109] In this embodiment, if there is no intermediate keyword in the intermediate keyword list that belongs to the preset instruction keyword library, it means that when the target user emits the target speech, the keyword corresponding to the blurry part of the target speech is not a keyword in the preset instruction keyword library. It may be that the pronunciation of the auxiliary word or modifier is blurry, resulting in a low matching degree between the target speech and the standard instruction in step S200. However, the only standard instruction that is matched is also correct.
[0110] Furthermore, the method may also include the following steps:
[0111] S800, if the number of target matches in the target match list is equal to 0, a preset voice error message is generated; wherein, the voice error message is used to prompt the target user to resend the voice message.
[0112] In this embodiment, if the number of target matching degrees in the target matching degree list is equal to 0, it means that no standard instruction that meets the requirements has been matched. At this time, the target user's pronunciation may be too unclear, resulting in the inability to recognize it. Therefore, a preset voice error prompt is generated to prompt the target user to re-speak.
[0113] In this embodiment, if the target user is in the driver's seat, several standard instructions are matched based on the semantic information of the target speech. If the number of matched labeled instructions is greater than 1, one of the matched labeled instructions is randomly determined as the target instruction corresponding to the target speech. At this time, the target instruction may be the actual instruction corresponding to the target speech or it may not be the actual instruction corresponding to the target speech, so that the target speech has a certain probability of matching the wrong standard instruction, so as to achieve the same situation as the co-pilot executing the wrong instruction during actual flight. Moreover, this method is not regular, and the target personnel cannot grasp the pattern of the wrong instruction, thereby improving the training effect.
[0114] Furthermore, if the number of matched labeled instructions is equal to 1, and the only target matching degree in the target matching degree list is greater than or equal to the second preset matching degree threshold, then the standard instruction corresponding to the only target matching degree in the target matching degree list is determined as the target instruction corresponding to the target speech; at this time, the matched target instruction is the actual instruction corresponding to the target speech, that is, the correct instruction; thereby ensuring that the target speech can match the correct standard instruction in most cases, so as to achieve the purpose of simulation training.
[0115] Furthermore, if the number of matched labeled instructions is equal to 1, and the only target matching degree in the target matching degree list is less than the second preset matching degree threshold, then the target speech is converted into the corresponding target text, and the target instruction corresponding to the target speech is determined based on the target text. At this time, the determined target instruction may be a correct instruction or an incorrect instruction, which further makes the simulation training conform to the actual scenario and improves the training effect.
[0116] Example 3:
[0117] Based on the above embodiment two, in a real flight scenario, the co-pilot executes the instructions issued by the pilot. As the flight duration increases and the flight status changes, the accuracy and efficiency of the co-pilot in executing the pilot's instructions will vary. When conducting flight training for users, to achieve optimal training efficiency, it is necessary to simulate real flight scenarios as much as possible. Therefore, the following method is provided:
[0118] Q100: If the target user is in the pilot's position, then acquire the target voice sent by the target user; wherein, the target user is any user conducting flight training.
[0119] In this embodiment, it is understood that when conducting simulated flight training for a target user, the target user can train as either the pilot or the co-pilot. If the target user is in the pilot's position, the target user is conducting simulated training as the pilot. Under normal circumstances, the pilot issues voice commands to the co-pilot, and the co-pilot executes the commands corresponding to the pilot's voice commands. When the target user is in the pilot's position and issues a voice command, the target voice command issued by the target user can be acquired. Whether the target user is in the pilot's position can be determined by image recognition or human detection sensors, which will not be elaborated here.
[0120] Q200, match the semantic information of the target speech with each preset standard instruction to obtain a standard matching degree list A = (A1, A2, ..., A...). i A n ), i=1, 2,..., n; among them, A i Let n be the matching degree between the target speech and the i-th standard instruction, and n be the number of preset standard instructions.
[0121] In this embodiment, several standard instructions are preset, such as turning left 30°, climbing up 1000m, etc. The target speech can be converted into text information first, and then the corresponding speech information is obtained. The standard instructions also have corresponding speech information. It should be noted that those skilled in the art can use existing semantic similarity determination methods according to actual needs to match the semantic information of the target speech with each preset standard instruction. This will not be elaborated here.
[0122] Q300, iterate through A, if A i If ≥α1, then A i Determine the target match degree; to obtain the target match list B = (B1, B2, ..., B...). j B m ), j = 1, 2, ..., m; where, B j The j-th target matching degree is determined, m is the number of target matching degrees determined; α1 is the preset first matching degree threshold.
[0123] In this embodiment, if A i ≥α1 indicates that the semantic information of the i-th standard instruction is quite similar to the speech information corresponding to the target speech, and the i-th standard instruction is likely to be the standard instruction corresponding to the target speech.
[0124] Furthermore, the value range of α1 is [0.7, 0.8]; for example: α1 = 0.8; α1 can be obtained based on experience or through a large number of experiments.
[0125] In this embodiment, α1 can be determined through the following steps:
[0126] Q310, Obtain the current training duration T of the target user. now Current ambient noise intensity J now Current flight altitude H now And the current flight speed V now .
[0127] In this embodiment, it is understood that as the simulation training time of the target user increases, the environmental noise intensity changes, the current flight altitude changes, and the current flight speed changes, the target user's state will be affected to a certain extent; for example, the longer the simulation training time, the more tired the target user becomes, and the slower the reaction speed will be.
[0128] Q320, according to T now J now H now and V now , determine α1.
[0129] Furthermore, step Q320 may include the following steps:
[0130] Q321, according to T now J now H now and V now The weights for training duration (ω1), environmental noise intensity (ω2), flight altitude (ω3), and flight speed (ω4) are determined; where ω1 = T now / TK;ω2=J now / JK;ω3=H now / HK;ω4=V now / VK; TK is the preset planned training duration; JK is the preset maximum ambient noise intensity; HK is the preset maximum flight altitude; VK is the preset maximum flight speed.
[0131] In this embodiment, T is processed through the above steps. now J now H now and V now Normalization is performed so that ω1, ω2, ω3 and ω4 are all in the range of 0 to 1.
[0132] Q322, based on ω1, ω2, ω3, and ω4, what is the target user's current fatigue level ρ? now = (ω1+ω2+ω3+ω4) / 4.
[0133] In this embodiment, ρ now Within the range of 0 to 1; T now J now H now and V now The larger the corresponding value, the greater the current fatigue level of the target user, the worse the target user's condition, and the lower the reaction ability.
[0134] Q323, obtain the preset fatigue level and matching degree threshold group list RH = (RH1, RH2, ..., RH2). e , ..., RH f ), e = 1, 2, ..., f; where RH e Here, f represents the preset fatigue level and matching threshold group, where f is the number of preset fatigue level and matching threshold groups; RH e =([RH e,1 RH e,2 ), RH e,3 ); [RH e,1 RH e,2 ) represents the preset fatigue level range for the e-th fatigue level; RH e,1 RH is the minimum value of the e-th fatigue level range. e,2 The maximum value within the range of the e-th fatigue level; RH e,3The matching threshold corresponding to the e-th fatigue level range; RH g,2 =RH g+1,1 , g = 1, 2, ..., f-1; RH g,3 <RH g+1,3 .
[0135] Q324, iterate through RH, if ρ now ∈[RH e,1 RH e,2 If ), then α1 = RH e,3 .
[0136] In this embodiment, it can be understood that ρ now The larger the value, the smaller the determined α1, which in turn makes m larger, meaning that the determined target matching degree is more, and correspondingly, the probability that the target command is not the actual command corresponding to the target voice is also greater, which is equivalent to the greater probability of the co-pilot control module malfunctioning.
[0137] By following the steps above, the flight status and training duration can be combined with the error probability of the co-pilot control module executing the command corresponding to the target voice, making the simulation training more in line with the actual flight scenario and improving the training effect.
[0138] Q400, if m>1, then randomly select the standard instruction corresponding to a target matching degree in B as the target instruction corresponding to the target speech; where NUM1 is the number of target matching degrees.
[0139] Understandably, normally a target speech corresponds to only one standard command. However, in this case, at least one standard command is identified, including both the actual command corresponding to the target speech and erroneous commands. Since the standard command corresponding to a target match degree in the target match degree list is randomly selected as the target command corresponding to the target speech, the identified target command may be incorrect. This simulates a real scenario where the co-pilot executes an incorrect command. Erroneous commands are generated randomly and without regularity, making it impossible for the target personnel to predict them in advance, thus improving the training effect.
[0140] Furthermore, after step Q400 and before step Q500, the method further includes the following steps:
[0141] Q410, if m=1, then the standard instruction corresponding to B1 is determined as the target instruction corresponding to the target speech.
[0142] In this embodiment, if m=1, it means that only one standard instruction is determined. Obviously, this standard instruction is the target instruction corresponding to the target speech.
[0143] Furthermore, after step Q400 and before step Q500, the method further includes the following steps:
[0144] Q420, if m=0, then generate a preset voice error message; wherein, the voice error message is used to prompt the target user to resend the voice message.
[0145] In this embodiment, if m=0, it means that the standard instruction could not be matched, and the target user may have abnormal pronunciation, so a preset voice error prompt is generated.
[0146] The Q500 executes the target command through a preset co-pilot control module.
[0147] Furthermore, after step Q500, the method further includes the following steps:
[0148] Q600: If the target command is different from the actual command corresponding to the target voice, determine whether the target user issues a correction command within a preset time period after the co-pilot control module has executed the target command.
[0149] Furthermore, step Q600 may include the following steps:
[0150] Q610, within a preset time period after the co-pilot control module completes the execution of the target command, acquire the actual flight state vector once at preset time intervals to obtain the actual flight state vector list PK = (PK1, PK2, ..., PK1). ε , ..., PK σ ), ε=1, 2,...,σ; where, PK ε σ represents the ε-th actual flight state vector obtained, where σ is the number of actual flight state vectors obtained.
[0151] In this embodiment, after the co-pilot control module executes the target command, the simulated flight state will continuously change. The actual flight state vector can be obtained once every preset time interval to obtain PK. The flight state vector may include flight altitude, flight speed, turn rate, climb rate, and descent rate, etc.
[0152] Q620, obtain each standard flight state vector after executing the actual command corresponding to the target voice, so as to obtain the standard flight state vector list PK' = (PK'1, PK'2, ..., PK') for PK. ε , ...,PK' σ ); where PK' ε The ε-th standard flight state vector after executing the actual command corresponding to the target voice.
[0153] If the actual command corresponding to the target voice is executed by the co-pilot control module, a standard flight state vector will be preset at preset time intervals within a preset time period after execution to obtain PK'.
[0154] Q630, obtain the similarity between each actual flight state vector in PK and the corresponding standard flight state vector in PK', to obtain a similarity list δ = (δ1, δ2, ..., δ... ε , …, δ σ ); where δ ε For PK ε With PK' ε The similarity between them.
[0155] Q640, if the similarity in δ decreases sequentially, it is determined that the target user did not issue a correction command within the preset time period after the co-pilot control module completed the target command.
[0156] In this embodiment, if the similarity in δ decreases sequentially, it indicates that the difference between the actual flight state and the corresponding standard flight state is getting bigger and bigger, and the target user has not issued a correction command.
[0157] Q650, if the similarity in δ increases first and then decreases sequentially, then it is determined that the target user issues a correction command within a preset time period after the co-pilot control module completes the target command.
[0158] In this embodiment, if the similarity in δ increases first and then decreases sequentially, it indicates that the difference between the actual flight state and the corresponding standard flight state increases first and then decreases, indicating that the target user issues a correction command.
[0159] Q700: If the target user does not issue a correction command within the preset time period TP after the co-pilot control module completes the target command, the target user is determined to be unqualified in training; otherwise, the time TU from the co-pilot control module completing the target command to the target user issuing the correction command is obtained.
[0160] Q800, based on TU, determine the target user's corrective response level θ = TU / TP.
[0161] The above steps can determine whether the target user's simulation training is qualified, and can also specifically quantify and evaluate the target user's corrective response.
[0162] In this embodiment, if the target user is in the driver's seat, several standard commands are matched based on the semantic information of the target speech. If the number of matched labeled commands is greater than 1, one of the matched labeled commands is randomly determined as the target command corresponding to the target speech. At this time, the target command may be the actual command corresponding to the target speech or it may not be the actual command corresponding to the target speech, so that the target speech has a certain probability of matching the wrong standard command, thus achieving the same situation as the co-pilot executing the wrong command during actual flight. In addition, the first matching degree threshold is determined based on the target user's current training time, current environmental noise intensity, current flight altitude, and current flight speed. The corresponding first matching degree threshold is different under different flight states, so the number of determined target matching degrees is also different, and thus the probability that the determined target command is the actual command corresponding to the target speech is also different. This makes the simulation training more in line with the actual scenario, and the method is not regular, so the target personnel cannot grasp the pattern of wrong commands, thereby improving the training effect.
[0163] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0164] Example 4:
[0165] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.
[0166] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0167] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0168] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0169] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0170] Example 5:
[0171] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0172] The electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments in this application.
[0173] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and a bus connecting different system components (including memory and processor).
[0174] The memory stores program code that can be executed by the processor, causing the processor to perform the steps in the various embodiments described in this specification.
[0175] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).
[0176] The memory may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0177] A bus can represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus structures.
[0178] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0179] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0180] Embodiments of the present invention also provide a computer program product including program code, which, when the program product is run on an electronic device, causes the electronic device to perform the steps of the methods described above in various exemplary embodiments of the present invention.
[0181] While specific embodiments of the invention have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention.
Claims
1. An intelligent flight training system, characterized in that, The system includes: a voice playback device, a storage medium, and a processor; wherein, the voice playback device is used to play virtual voice corresponding to the driver; the storage medium stores at least one instruction or at least one program segment, which is loaded and executed by the processor to implement the following steps: H100, if the target user is in the passenger seat, then obtain the time taken for each virtual voice command executed by the target user within the preset sliding time window TH, to obtain a time list T = (T1, T2, ..., T...). x ,…,T y ), x=1,2,…,y; where, T x The time taken to execute the instruction corresponding to the xth virtual voice in TH for the target user, where y is the number of virtual voice instructions executed by the target user in TH; H200, based on T, determine the average time T' = (1 / y) × ∑ for the target user to execute the virtual voice command within TH. y x=1 T x ; H300, if T' < T1 or T' > T2, then obtain the next virtual voice file YA to be played; where T1 is the first preset time consumption threshold, T2 is the second preset time consumption threshold; T1 < T2; H400, if there is a preset reverse instruction for the instruction corresponding to YA, then the first preset virtual voice processing method is used to process YA to obtain the reverse virtual voice file YB corresponding to YA. H500 plays YB through the voice playback device and obtains the actual command YB' corresponding to the target user's actual operation on YB; H600 determines the target user's completion rate for YA based on the instructions corresponding to YB' and YA; Step H400 includes the following steps: H410, retrieve the instruction ZA corresponding to YA; H420, retrieve the inverse instruction ZA' corresponding to ZA; H430, extract keywords from ZA and ZA' to obtain the keyword list GA=(GA1, GA2, ..., GA') corresponding to ZA. u , ...,GA v ) and the keyword list corresponding to ZA' GA' = (GA'1, GA'2, ..., GA') u , ...,GA' v ),u=1,2,…,v; among them, GA u For the u-th keyword corresponding to ZA, GA' u Let be the u-th keyword corresponding to ZA', and v be the number of keywords corresponding to ZA and ZA'; H440, iterate through GA and GA', if GA u with GA' u If they are different, then GA' u This has been identified as the reverse keyword corresponding to ZA'; H450, adjust the volume of the audio file segment corresponding to the reverse keyword to the first volume YL1 to obtain YB; where YL1=λ×YL; YL is the volume preset by the target user; λ is the preset reverse keyword volume weight, λ<1.
2. The intelligent flight training system according to claim 1, characterized in that, After step H400 and before step H500, the at least one instruction or the at least one program segment is loaded and executed by the processor to perform the following steps: H460, if there is no preset reverse instruction for the instruction corresponding to YA, then the second preset virtual voice processing method is used to process YA to obtain the reverse virtual voice file YB corresponding to YA.
3. The intelligent flight training system according to claim 2, characterized in that, Step H460 includes the following steps: H461, retrieve the volume of each preset for the target user within a preset historical time period to obtain a historical preset volume list HE = (HE1, HE2, ..., HE...). a HE b ), a=1,2,…,b; where, HE a b represents the a-th preset volume for the target user within a preset historical time period, and b represents the number of preset volumes for the target user within the preset historical time period. H462 uses a preset clustering algorithm to cluster the volume within HE to obtain several clusters; H463, the average volume corresponding to the cluster with the most volume is determined as the target volume ML corresponding to YA; H464, adjust the volume of YA to the second volume YL2 to obtain YB; where YL2=η×ML; η is the preset volume adjustment weight; η<1.
4. The intelligent flight training system according to claim 1, characterized in that, Step H600 includes the following steps: H610, if the instructions corresponding to YB' and YA are the same, then obtain the duration TG from the end of YB playback to the target user completing YB'; H620, if TG∈[TG1,TG2], then determine the target user's completion degree DQ for YA = (TG1 / 2 + TG2 / 2) / TG; where TG1 is the preset minimum instruction completion time and TG2 is the preset maximum instruction completion time; H630, if TG < TG1 or TG > TG2, then the target user's completion rate DQ for YA is determined to be 0.
5. The intelligent flight training system according to claim 1, characterized in that, After step H600, the at least one instruction or the at least one program segment is loaded and executed by the processor, and the following steps are also performed: H700, obtain the completion rate of the target user's command for each virtual voice, to obtain a completion rate list FE = (FE1, FE2, ..., FE2). c , ..., FE d ), c=1,2,…,d; where, FE c Let d represent the completion rate of the command given by the target user to the c-th virtual voice, and d represent the number of virtual voices. H710, based on FE, determine the average completion rate for the target user: FE' = (1 / d)∑ d c=1 FE c ; H720, if FE'≥FR, then the target user is deemed to have passed the training; otherwise, the target user is deemed to have failed the training.
6. The intelligent flight training system according to claim 1, characterized in that, The value of λ ranges from 0.2 to 0.
5.
7. The intelligent flight training system according to claim 3, characterized in that, The value of η ranges from 0.3 to 0.4.
Citation Information
Patent Citations
Experiential flight simulation training system and simulation training method
CN114067631A
Virtual-real combined immersive driving simulation training simulation system
CN116580619A