A multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-08-14
AI Technical Summary
然而,高原智慧舱调控参数时,制氧机、变频增压泵等设备联动运行,产生的机械噪声与气流扰动会形成渐进式的隐性干扰,此类渐进式的隐性干扰不会使语音显性指标、识别置信度明显下跌,仅在正常范围小幅波动,未达到云端切换阈值,却会持续降低本地语音识别的稳定性与准确性,导致系统无法预判识别劣化趋势,无法及时切换至云端抗干扰识别,影响智慧舱使用体验与运行安全性;
Smart Images

Figure CN122575365A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction technology, specifically to a multi-parameter voice control system for a high-altitude smart cabin that integrates voice interaction. Background Technology
[0002] In high-altitude engineering construction, scientific research, and border defense, the high-altitude smart cabin serves as the core enclosed space for personnel. The precise control of key environmental parameters such as oxygen concentration, air pressure, temperature, and humidity within the cabin directly impacts the safety of the personnel. Due to the extreme environment of low oxygen and low pressure at high altitudes, personnel are prone to altitude sickness and operational difficulties. Traditional contact-based control methods are insufficient to meet the needs for convenient and safe control. Therefore, contactless voice interaction technology is widely used in the voice control system of the high-altitude smart cabin. Currently, most high-altitude smart cabin voice control systems adopt an offline-first, online-second working mode. The core purpose is to adapt to the unreliable network conditions in high-altitude areas. In scenarios such as uninhabited high-altitude areas and border camps, there are often problems such as missing mobile signals, satellite network lag or interruption. If voice control relies too much on online recognition, network anomalies will directly lead to the paralysis of core control functions and cause safety hazards. Therefore, existing systems default to placing the voice interaction module in offline mode and do not automatically switch with network connectivity. They only manually or semi-automatically switch to online mode when non-core auxiliary needs are recognized or when the reliability of local voice recognition is insufficient. Among them, scenarios where non-core auxiliary needs trigger online switching include: when users issue non-life-saving commands such as "turn on online voice" or "upload control logs", the system connects to the cloud and uses the cloud's large model to complete functions such as complex semantic understanding and data uploading. Switching when local voice recognition is unreliable is based on the confidence level of the local recognition output: there is no need to fully parse the voice command, only monitor the confidence level index. When external linear indicators such as voice clarity and signal-to-noise ratio are abnormal, the confidence level will drop below the preset threshold, which will determine that the recognition is unreliable and trigger online switching, uploading the command to the cloud for parsing. However, when adjusting parameters in the high-altitude smart cabin, the oxygen generator, variable frequency booster pump, and other equipment operate in tandem, and the resulting mechanical noise and airflow disturbances create a gradual, implicit interference. This gradual, implicit interference does not cause a significant drop in explicit voice indicators or recognition confidence; it only fluctuates slightly within the normal range and does not reach the cloud switching threshold. However, it continuously reduces the stability and accuracy of local voice recognition, causing the system to be unable to predict the trend of recognition degradation and unable to switch to cloud-based anti-interference recognition in a timely manner, thus affecting the user experience and operational safety of the smart cabin. To address the above problems, this invention proposes a solution. Summary of the Invention
[0003] The purpose of this invention is to provide a multi-parameter voice control system for a high-altitude smart cabin that integrates voice interaction, in order to solve the problems mentioned in the background art.
[0004] This invention provides a multi-parameter voice control system for a high-altitude intelligent cabin that integrates voice interaction, comprising: The environmental monitoring module is used to collect environmental monitoring data of the target cabin in real time. The voice dispatch module is used to determine whether the voice command is a local or cloud-based recognition subject based on the network connectivity status of the target cabin when the voice command is collected, according to the voice command issued by the personnel stationed in the target cabin. The network connectivity status is either online or offline. If the target cabin's network connectivity is in a connected state, a preliminary determination is made of the entity that recognizes and executes the voice command. Based on the preliminary determination result, it is determined whether the entity that recognizes and executes the voice command is a local entity. If it is not a local entity, an advanced determination is made of the voice command. Based on the advanced determination result, it is determined whether the entity that recognizes and executes the voice command is a local entity or a cloud entity. If the network connectivity status of the target cabin is no network, then the recognition of the voice command is directly determined to be local recognition. The data analysis module is used to collect and generate a fixed amount of analytical data that can be analyzed. The content of any analytical data collected and generated is as follows: when any voice command is detected, if the voice command meets the preset collection conditions, the analytical data of the voice command is generated. The analytical data includes the first data of the voice command obtained from the cloud recognition module, the second data of the voice command obtained from the environmental monitoring module, and the third data of the voice command obtained from the voice scheduling module. The data analysis module is also used to analyze all the generated analysis data after the amount of collected and generated analysis data reaches a preset fixed amount, and generate matching data for several preloaded sets. The cloud recognition module is used to input each voice command received by the cloud recognition execution entity into the pre-trained cloud big model for recognition and parsing. The cloud big model outputs the corresponding voice text. After the recognition is completed, it automatically records and stores all the processing algorithms used by the cloud big model to recognize the voice command. The local recognition module is used to input each received voice command, which is performed by a local recognition entity, into a pre-trained local large model to recognize and output the corresponding speech text.
[0005] Furthermore, the initial determination of the entity responsible for recognizing the voice command is as follows: The PESQ method is used to output the confidence level of the voice command. Starting from the time the voice command is issued, a preset comparison duration P4 is traced backward to define the comparison period. All voice commands issued by personnel stationed in the target cabin within this comparison period are obtained. The confidence level difference between the two voice commands with the closest issuance times is calculated sequentially. If any calculated confidence level difference is greater than or equal to a preset confidence level difference threshold, an advanced determination is performed on the recognition of the voice command. If all calculated confidence level differences are less than the preset confidence level difference threshold, the entity responsible for recognizing the voice command is determined to be a local recognition entity.
[0006] Furthermore, the preset data collection conditions are as follows: Starting from the moment the voice command is issued, a preset environmental change observation duration P1 is traced back to determine the corresponding end time, thereby defining the collection period. All voice commands issued within the collection period are acquired. If the recognition execution entity of all voice commands is local recognition, and the confidence of all voice commands shows a continuous downward trend on the line graph, and the voice command is a correction of the previous voice command, then the voice command is determined to meet the preset collection conditions. Here, the voice command is a correction of the previous voice command, which is a subsequent voice command issued by the stationed personnel to correct the recognition abnormality caused by the inaccurate recognition result of the previous voice command.
[0007] Furthermore, the content of the matching data for generating several preloaded sets is as follows: S11: Label all generated analysis data as A1, A2, ..., Aa, where a≥1; label all environmental parameters selected by the management personnel as B1, B2, ..., Bb, where b≥1. S12: According to the preset judgment steps, determine whether the environmental parameters B1, B2, ..., Bb are interference parameters of the analysis data A1, and re-mark all interference parameters that are judged to be analysis data A1 as D1, D2, ..., Dd, 1≤d≤b; S13: Select several voice command issuance times from the third data in the analysis data A1 as gradient inflection points of the analysis data A1. Combine the acquisition times corresponding to the monitoring values that are farthest and closest to the current time in the analysis data A1 to determine the interference intervals G1, G2, ..., Gg, where 1≤g≤f+1, and f is the total number of gradient inflection points of the selected analysis data A1. S14: Calculate the average interference weights of the interference parameters D1, D2, ..., Dd within the analysis data A1; S15: Calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data in the analysis data A1; S16: Calculate the interference weights of all interference parameters in the analysis data A2, A3, ..., Aa according to S14; calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data in the analysis data A2, A3, ..., Aa according to S15. S17: Based on the first data in the analysis data A1, A2, ..., Aa, construct several preloaded sets, labeled as M1, M2, ..., Mm, where 1≤m≤a; S18: Generate matching data for preloaded sets M1, M2, ..., Mm according to preset generation steps, and transfer the matching data of preloaded sets M1, M2, ..., Mm to the voice scheduling module for storage. The generation steps are as follows: SS51: Extract all analysis data from analysis data A1, A2, ..., Aa that match the preloaded set M1, and relabel them as N1, N2, ..., Nn, where 1 ≤ n <a; SS52: Obtain all interference parameters of the analysis data N1, add all the interference parameters to an empty set to obtain the interference set of the analysis data N1, and similarly obtain the interference sets of the analysis data N2, N3, ..., Nn respectively. Iterate through all the obtained interference sets and select the interference set with the most consistent values as the associated set of the preloaded set M1. SS53: Relabel all analytical data within the analytical data N1, N2, ..., Nn, whose interference set is the association set, as Q1, Q2, ..., Qq, 1≤q≤n; SS54: For each interference parameter in the association set, obtain the average interference weight of the interference parameter in the analysis data Q1, Q2, ..., Qq respectively, and calculate its mean using the summation and averaging formula. Use the mean as the standard interference weight of the interference parameter in the association set of the preloaded set M1. SS55: Calculate the mean value of the interference coefficient of the first data in the analysis data Q1, Q2, ..., Qq using the summation and averaging formula, and use the mean value as the mean interference coefficient of the preloaded set M1. Similarly, calculate the mean value of the confidence coefficient and the mean value of the duration coefficient of the first data in the analysis data Q1, Q2, ..., Qq, and use the mean value as the mean confidence coefficient and the mean value of the duration coefficient of the preloaded set M1, respectively. SS56: Generate matching data for preloaded set M1 based on the average interference coefficient, average confidence coefficient, average duration coefficient, and the standard interference weights of all interference parameters in the associated set of preloaded set M1. SS57: Calculate and obtain matching data for the preloaded sets M2, M3, ..., Mm in sequence according to SS51 to SS56.
[0008] Furthermore, in S12, the determination steps are as follows: SS11: Extract all monitoring values of environmental parameter B1 from the second data set of analysis data A1, and label all extracted monitoring values as C1, C2, ..., Cc, where c≥1, according to the order of their acquisition time from the current time to the earliest. SS12: Determine whether environmental parameter B1 is an interference parameter of analysis data A1. The determination is as follows: Use a discrete point filtering algorithm to process the monitoring values C1, C2, ..., Cc, and map all the remaining monitoring values after data processing onto a line graph. If all monitored values on the line graph show a decreasing or increasing trend from left to right, then environmental parameter B1 is determined to be an interference parameter of analytical data A1; otherwise, environmental parameter B1 is determined not to be an interference parameter of analytical data A1. SS13: Determine whether environmental parameters B2, B3, ..., Bb are interference parameters of analysis data A1 in SS11 to SS12. After the determination, re-label all interference parameters determined to be analysis data A1 as D1, D2, ..., Dd, where 1≤d≤b.
[0009] Furthermore, S13, the specific content is as follows: SS21: Extract the confidence scores of all voice commands from the third data in the analysis data A1, and label all extracted confidence scores as E1, E2, ..., Ee in order of their distance from the current time, from farthest to closest, where e≥1; SS22: Compare the confidence scores E1 and E2. If the confidence score E1 is greater than E2 and the value of confidence score E1-E2 is greater than or equal to the preset standard gradient difference, then select the time of speech command issuance corresponding to confidence score E2 as a gradient inflection point; otherwise, no processing is performed. SS23: Following the steps in SS22, compare the confidence levels E2 and E3, E3 and E4, ..., Ee-1 and Ee in sequence. After the comparison, relabel all gradient inflection points in the order they were selected, from first to last, as F1, F2, ..., Ff, where 1≤f <e; SS24: Traverse the second data in the analysis data A1, mark the collection time corresponding to the monitoring value farthest from the current time as the starting point, and mark the collection time corresponding to the monitoring value closest to the current time as the ending point. Combine the gradient inflection points F1, F2, ..., Ff to divide into several interference intervals, and mark them in chronological order from first to last as G1, G2, ..., Gg, 1≤g≤f+1. Among them, interference interval G1 corresponds to the time from the starting point to the gradient inflection point F1, G2 corresponds to the time from the next time of gradient inflection point F1 to gradient inflection points F2, G3, ..., Gg, and so on.
[0010] Furthermore, in S14, the calculation steps for the average weight of the interference parameters D1, D2, ..., Dd within the analysis data A1 are as follows: SS31: Extract the maximum and minimum monitoring values from all monitoring values of interference parameter D1 in the second data, calculate the difference between the two extracted monitoring values, and take the difference as the parameter variation amplitude H1-1 of interference parameter D1 in interference interval G1. Similarly, calculate the parameter variation amplitude H1-2, H1-3, ..., H1-g of interference parameter D1 in interference intervals G2, G3, ..., Gg. SS32: Obtain the confidence scores of all voice commands issued within the interference interval G1 from the third data, extract the highest and lowest confidence scores, and calculate the difference to obtain the confidence score variation amplitude I1 within the interference interval G1. SS33: Utilizing Formulas Calculate the interference percentage J1 of interference parameter D1 in interference interval G1, where Hi-1 represents each of the parameter variation amplitudes H1-1, H1-2, ..., H1-g; SS34: Calculate and obtain the interference percentages J2, J3, ..., Jg of interference parameter D1 in the interference intervals G2, G3, ..., Gg in sequence according to SS33; SS35: A discrete point filtering algorithm is used to process J1, J2×η1, J3×η2, ..., Jg×ηg-1 to obtain the average value of all remaining data after data processing. The average value is then calibrated as the average interference weight of interference parameter D1, where η1, η2, ..., ηg-1 are preset compensation coefficients for interference parameter D1 as the value of the marked index of the interference interval increases. SS36: Calculate and obtain the average interference weights of interference parameters D2, D3, ..., Dd in sequence according to SS31 to SS35.
[0011] Furthermore, S15, the calculation is as follows: SS41: Constructing an equation based on the interference interval G1 In the equation, a1, a2, and a3 are the interference parameter coefficients, confidence level variation coefficients, and interval duration coefficients to be solved, respectively; L1 is the interval duration of the interference interval G1; Hj-1 refers to each of the parameter variation amplitudes H1-1, H2-1, ..., Hd-1; H2-1, ..., Hd-1 are the parameter variation amplitudes of the interference parameters D2, D3, ..., Dd in the interference interval G1, calculated when calculating the average interference weights of the interference parameters D2, D3, ..., Dd. SS42: Construct equations sequentially based on interference intervals G2, G3, ..., Gf+1, following the pattern in SS41. The equation constructed based on interference interval G2 is... In the formula, β1, β2, and β3 are the preset first, second, and third compensation coefficients, respectively; SS43: By constructing a system of three linear equations in three variables based on any three equations of the interference intervals G1, G2, ..., Gg, we can obtain C. g 3 Solve each of the three systems of linear equations in three variables to obtain C. g 3 The values of a1, a2 and a3; SS44: Using a discrete point filtering algorithm for C f+1 3 The values of a1 are processed to obtain the average value of all remaining a1 values after data processing. The average value is then calibrated as the interference coefficient of the first data in the analysis data A1. SS45: Apply discrete point filtering algorithms to C according to SS43 to SS44 respectively. f+1 3 The values of a2 and a3 are processed to obtain the average value of all remaining a2 and a3 values after data processing. The average value is then labeled as the confidence coefficient and duration coefficient of the first data in the analysis data A1.
[0012] Furthermore, the advanced determination of the voice command is as follows: Starting from the moment the voice command is issued, a preset observation duration P3 is traced back to define the observation period. Based on the observation period, all interference parameters relative to the voice command are selected from all environmental parameters selected by the management personnel. All interference parameters within the observation period are integrated into an interference set. The interference set is used as a search condition. If an associated set that matches the interference set is found, the matching data of the preloaded set corresponding to the associated set is obtained. Otherwise, it is determined that the recognition execution subject of the voice command is local recognition. The interference mean coefficient of each interference parameter is obtained from the matching data, and the formula is used. Calculate and obtain the predicted recognition duration T1 based on the voice command, where Uu refers to each interference parameter in the observation period, λu represents the interference standard weight of the corresponding interference parameter from the matching data, X1, X2, and X3 are the average interference coefficient, average confidence coefficient, and average duration coefficient contained in the matching data, respectively, and v is the total number of interference parameters determined to be interference parameters in the observation period. If T1≈0, the entity executing the voice command recognition is determined to be cloud-based recognition; otherwise, the entity executing the voice command recognition is determined to be local recognition. At this time, a task timing instruction is generated based on the calculated predicted recognition duration T1, and the task timer is started. The task timer is initially assigned a value of 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1, the entity executing the recognition of all voice commands collected subsequently is automatically determined to be cloud-based recognition, and no network status or any subsequent determinations are made. The pre-fetching recognition time T1 and the pre-loaded set are simultaneously transmitted to the cloud recognition module.
[0013] Furthermore, after receiving the pre-fetch recognition duration T1 and the preloaded set, the cloud recognition module synchronously starts the task timer. The task timer is initially assigned a value of 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1×ε1, the preloading operation is automatically executed to preload all the processing algorithms contained in the preloaded set into the runtime cache.
[0014] Compared with existing technologies, it has the following advantages: This invention collects real-time environmental monitoring data within the target cabin using an environmental monitoring module and collects voice commands issued from within the target cabin using a voice dispatch module. It determines the execution entity based on the network connectivity status within the target cabin and selects whether to perform voice command recognition locally or in the cloud based on the execution entity. This method achieves intelligent adaptive switching between local and cloud-based recognition, ensuring stable operation of core voice control functions in network-free scenarios and avoiding control paralysis and security risks caused by network interruptions. Furthermore, when the network is connected, it fully utilizes the high-precision recognition advantages of the cloud-based large model, compensating for the shortcomings of the local model in complex semantic understanding and anti-interference capabilities, thus improving the flexibility and practicality of voice control in the high-altitude smart cabin. This invention collects a fixed amount of analytical data through a data acquisition module. The analytical data shows a continuous downward trend in the confidence level of voice commands recognized in the cloud over a certain period of time on a line graph. All collected analytical data are analyzed to determine the matching data of several preloaded sets. In the determination process, the changes in the monitored values of each interference parameter under each abnormal voice command in each interference interval are analyzed, as well as the interference ratio of each interference parameter in the change of confidence level of the voice command. Combined with all the processing algorithms used in the cloud recognition process of each abnormal voice command, several preloaded sets are obtained based on the activated processing algorithms. The interference coefficient of each preloaded set is calculated based on the interference ratio and the change in confidence level. The association set, average interference coefficient, average confidence coefficient, and average duration coefficient of each preloaded set are determined by combining the influence of each interference parameter on confidence level and the duration of abnormal recognition. In this way, the problem of existing technologies being unable to identify such small fluctuation interference and unable to predict the occurrence of recognition degradation trends is solved. The problem of recognition errors caused by implicit interference is reduced from the root, and the stability and accuracy of speech recognition are improved. This invention defines an observation period for voice commands that perform advanced judgments, analyzes and determines all interference parameters relative to the voice command within that observation period, matches the corresponding association sets based on all interference parameters, finds the corresponding preloaded set, and calculates the average interference coefficient, average confidence coefficient, and average duration coefficient based on the association sets, and calculates the observation and recognition duration. Voice commands collected after the observation and recognition duration are uniformly recognized by the cloud. In this way, for the gradual implicit interference generated by the linkage of equipment in the high-altitude cabin, the trend of interference evolution can be determined in advance without waiting for the confidence to drop below the threshold, and the cloud recognition mode can be switched in time. This effectively avoids the problems of decreased recognition accuracy and frequent command correction caused by the continuous impact of implicit interference, and further improves the timeliness and reliability of voice control. This invention determines the processing algorithms that should be activated for a voice command after the observed recognition time by observing the recognition time, and preloads these processing algorithms in advance. This method effectively solves the technical problem of algorithm loading delay during cloud recognition. By preloading the corresponding processing algorithms into the runtime cache, when subsequent voice commands switch to cloud recognition, the preloaded algorithms can be directly called without reloading, which greatly shortens the response time of voice recognition, avoids the control lag caused by slow algorithm loading, ensures the efficiency and smoothness of cloud recognition, and further improves the experience of voice interaction and in-cabin control. This invention effectively utilizes network resources when a network is available, further optimizing system operating efficiency and recognition performance. Specifically, the system does not blindly switch to cloud recognition as soon as a network connection is established. Instead, it uses a layered logic of initial and advanced judgments, combined with multi-dimensional indicators such as voice command confidence, timing difference, and environmental interference parameters, to determine the recognition execution subject. It only switches to cloud recognition when local recognition shows a deterioration trend or significant interference, thus avoiding waste of network resources. At the same time, through pre-loading set matching and algorithm pre-loading, cloud computing resources and local resources are rationally allocated. While fully leveraging the strong anti-interference advantages of the large cloud model, unnecessary cloud requests are avoided, reducing network data transmission load. This adapts to the limited network resources in high-altitude areas, achieving optimal matching between network resources and system performance, balancing recognition accuracy, response speed, and resource utilization. Attached Figure Description
[0015] Figure 1 This is a system block diagram of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Please see Figure 1 This application provides a multi-parameter voice control system for a high-altitude smart cabin that integrates voice interaction, including an environmental monitoring module, a voice dispatching module, a local recognition module, a cloud recognition module, and a data analysis module; The environmental monitoring module is used to collect environmental monitoring data of the target cabin in real time, wherein the environmental monitoring data includes real-time monitoring values of several environmental parameters. In this application, the target cabin refers to the plateau smart cabin. The environmental parameters are selected by the management personnel after screening, taking into account the actual application scenario of the plateau smart cabin, the life safety needs of the stationed personnel, the system voice control target and the voice recognition anti-interference requirements. These parameters include, but are not limited to, oxygen concentration, cabin air pressure, ambient temperature, relative humidity, equipment operating noise intensity and harmful gas concentration. The voice dispatch module is used to collect voice commands issued by any resident personnel in the target cabin and simultaneously record the time when the voice command is issued. The voice dispatch module is also used to determine whether the recognition execution entity of a voice command is local recognition or cloud recognition, based on the network connectivity status of the target cabin at the time of acquisition, for each voice command acquired. If the target cabin is in a network connectivity state, a preliminary determination is made regarding the entity responsible for recognizing and executing the voice command. The preliminary determination is as follows: The PESQ method is used to output the confidence level of the voice command. Starting from the time the voice command is issued, a preset comparison time P4 is traced back to define the comparison period. All voice commands issued by the personnel stationed in the target cabin within the comparison period are obtained. The confidence level difference between the two voice commands with the closest issuance times is calculated sequentially. If any confidence level difference is greater than or equal to a preset confidence level difference threshold, an advanced determination is performed on the entity that recognizes and executes the voice command. If all calculated confidence level differences are less than the preset confidence level difference threshold, the entity that recognizes and executes the voice command is determined to be a local entity. Among them, the PESQ method is an objective speech quality assessment method based on ITU-T Recommendation P.862. Its core function is to output a confidence score or quality score of 0–4.5 based solely on the acoustic characteristics of speech, such as environmental noise, timbre, distortion, proficiency, and signal-to-noise ratio, without considering the text content. The following is an example of calculating the confidence difference between the two voice commands whose issuance times are closest to each other among all the acquired voice commands: If there are 5 voice commands Z1, Z2, Z3, Z4, and Z5, and their issuance times are progressively further away from the current time, then the differences to be calculated are Z1-Z2, Z2-Z3, Z3-Z4, and Z4-Z5. The advanced judgment criteria are as follows: Starting from the moment the voice command is issued, a preset observation period P3 is traced back to define the observation period. The value of P3 is based on 1 / 2 to 2 / 3 of P1, which can cover the gradual change of environmental parameters in the recent period and avoid data redundancy caused by excessively long backtracking time, thus ensuring prediction efficiency. At the same time, the confidence level of the current voice command and the changes in environmental parameters during the observation period are recorded as the basis for predicting the interference of the next voice command. From all environmental parameters selected by the administrator, all interference parameters relative to the voice command during the observation period are filtered out, as follows: For each environmental parameter, all monitoring values of the environmental parameter within the observation period are obtained. A discrete point filtering algorithm is used to process all the monitoring values. All remaining monitoring values after data processing are mapped onto a line graph. In the line graph, the horizontal axis represents the acquisition time and the vertical axis represents the monitoring value. All remaining monitoring values after data processing are mapped onto the line graph in order from the acquisition time to the current time, from farthest to closest. If all the monitoring values on the line graph show a sequential decrease or increase from left to right, the environmental parameter is determined to be an interference parameter during the observation period. Otherwise, it is determined not to be an interference parameter during the observation period. All interference parameters within the observation period are integrated into an interference set. The interference set is used as a search condition. If an associated set that matches the interference set is found, the matching data of the preloaded set corresponding to the associated set is obtained. Otherwise, the recognition execution subject of the voice command is determined to be local recognition. Here, "matching" means that the processing algorithms included are completely identical, excluding the order of the processing algorithms in the set. The interference mean coefficient of each interference parameter is obtained from the matching data, and the formula is used. Calculate and obtain the predicted recognition duration T1 based on the voice command, where Uu refers to each interference parameter in the observation period, λu represents the interference standard weight of the corresponding interference parameter from the matching data, X1, X2, and X3 are the average interference coefficient, average confidence coefficient, and average duration coefficient contained in the matching data, respectively, and v is the total number of interference parameters determined to be interference parameters in the observation period. If T1≈0, then the entity that recognizes and executes the voice command is determined to be cloud-based. Conversely, if the voice command is not recognized, the recognition execution entity is determined to be local recognition. At this time, a task timing instruction is generated based on the calculated predicted recognition duration T1 and the task timer is started. The task timer is initially assigned a value of 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1, the recognition execution entity of all voice commands collected subsequently is automatically determined to be cloud recognition, and no network status or any subsequent determinations are performed. The pre-fetch recognition duration T1 and the pre-loaded set are simultaneously transmitted to the cloud recognition module; It should be noted that the task timer has an automatic reset mechanism, and will be reset to zero only when any of the following conditions are met: When the administrator performs a manual reset operation, the key operating parameters of the equipment inside the target cabin change, and the system does not collect any voice commands within the set silence period; The specific value of ε1 needs to be determined by combining the number of processing algorithms, the type of algorithms, and the complexity of the algorithms contained in the preload set, to ensure that the preload operation can be completed before the task timer reaches T1, so as to provide efficient support for the cloud recognition of subsequent voice commands and avoid the recognition response speed being affected by the algorithm loading delay. The voice dispatch module is also used to transmit voice commands that are determined to be executed by local recognition to the local recognition module, and voice commands that are determined to be executed by cloud recognition to the cloud recognition module. The cloud-based recognition module is used to input the received voice command into the pre-trained cloud-based big data model for recognition and parsing. The cloud-based big data model outputs the corresponding voice text. After the recognition is completed, the system automatically records and stores all the processing algorithms used by the cloud-based big data model to recognize the voice command. Then, based on the voice text content, it is parsed into the corresponding equipment control command or voice question and answer response content and transmitted to the target cabin system, where the system executes the corresponding feedback and control operations. After receiving the pre-fetch recognition duration T1 and the preloaded set, the cloud recognition module synchronously starts the task timer. The task timer is initially assigned a value of 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1×ε1, the preloading operation is automatically executed to preload all the processing algorithms contained in the preloaded set into the runtime cache. If the target cabin is in a state of no network, the recognition and execution of the voice command is directly determined to be the local recognition, and the voice command is transmitted to the local recognition module. Due to the constraints of hardware computing power and model size, the local large model is weaker than the cloud large model in terms of recognition accuracy, complex semantic understanding and generalization ability. It is mainly used for basic voice recognition guarantee in the case of network outage. The local recognition module is used to input the received voice command into the pre-trained local large model to recognize and output the voice text, and then convert the voice text into control commands or local question and answer information that the system can execute to complete the corresponding voice interaction and cabin control actions. The data analysis module is used to collect and generate a fixed amount of analytical data for analysis. The fixed amount is set so that the amount of analytical data collected and generated is sufficient to fully cover the changing characteristics of environmental parameters, and the amount of data meets the validity and representativeness required for subsequent analysis, ensuring that the analysis results are true and reliable. The content of any analysis data collected is as follows: When any voice command is detected, if the voice command meets the preset collection conditions, analysis data of the voice command is generated. The analysis data includes first data of the voice command obtained from the cloud recognition module, second data of the voice command obtained from the environmental monitoring module, and third data of the voice command obtained from the voice scheduling module. The preset data collection conditions are as follows: Starting from the moment the voice command is issued, the preset environmental change observation duration P1 is traced back to determine the corresponding end time, thereby defining the collection period. All voice commands issued within the collection period are obtained. If the recognition execution entity of all voice commands is local recognition, and the confidence of all voice commands shows a continuous downward trend on the line graph, and the voice command is a correction of the previous voice command, then the voice command is determined to meet the preset collection conditions. The value of P1 is based on: covering the slow environmental change cycle caused by the linkage operation of equipment in the target cabin, ensuring that the condition determination can be completed within the range of gradual evolution of interference. The voice command is a correction to the previous voice command. In response to the situation where the previous voice command has an abnormal recognition result due to inaccurate recognition, the stationed personnel issue a subsequent voice command to correct the recognition abnormality. The previous voice command refers to the voice command that is closest to the time when the previous voice command was issued. It should be noted that the number of voice commands issued during the collection period is at least the minimum number of valid samples P2. The value of P2 is determined by limiting the number of valid voice commands during the collection period to no less than P2, ensuring that the confidence decline trend has statistical validity, and avoiding misjudgment due to insufficient samples. It should be noted that the determination method for the continuous downward trend in the confidence level of all voice commands on the line graph is as follows: A discrete point filtering algorithm is used to process the confidence scores of all voice commands. The remaining confidence scores after data processing are then mapped onto a line graph. In the line graph, the horizontal axis represents the time of issuance and the vertical axis represents the confidence score. All confidence scores after data processing are mapped onto the line graph in order from the time of issuance to the time of current issuance, from farthest to closest. If the confidence scores decrease sequentially, it indicates that the confidence scores of all voice commands show a continuous downward trend on the line graph. In this application, the discrete point filtering algorithm is the H-score filtering algorithm, which is used to filter out abnormal discrete points in the confidence score to ensure the accuracy of trend determination; The content of the first data obtained from the cloud recognition module for the voice command is as follows: The cloud recognition module retrieves all algorithm types used by the cloud recognition model when recognizing the voice command, and filters and extracts all processing algorithms used to eliminate interference in the voice command. The first data of the voice command is obtained based on all the extracted processing algorithms. In this application, the types of processing algorithms include, but are not limited to, low-frequency noise reduction, airflow filtering, speech enhancement, signal-to-noise ratio restoration, steady-state interference suppression, and adaptive noise cancellation; The second data obtained from the environmental monitoring module for the voice command is as follows: The second data of the voice command is generated by extracting all environmental monitoring data collected during the collection period from the environmental monitoring module and generating the second data based on the data. The second data also includes environmental monitoring data collected at the time the voice command is issued. The third data obtained from the voice dispatch module for the voice command is as follows: The confidence scores of all voice commands issued during the collection period are extracted from the voice scheduling module, and third data of the voice commands are generated based on them. The third data also includes the confidence scores of the voice commands. The data analysis module is also used to analyze all the generated analysis data after the amount of collected and generated analysis data reaches a preset fixed amount. The analysis content is as follows: S11: Label all generated analysis data as A1, A2, ..., Aa, where a≥1; label all environmental parameters selected by the management personnel as B1, B2, ..., Bb, where b≥1. S12: Following the preset judgment steps, determine whether environmental parameters B1, B2, ..., Bb are interference parameters of analysis data A1, and re-label all interference parameters determined to be interference parameters of analysis data A1 as D1, D2, ..., Dd, where 1 ≤ d ≤ b. The judgment steps are as follows: SS11: Extract all monitoring values of environmental parameter B1 from the second data set of analysis data A1, and label all extracted monitoring values as C1, C2, ..., Cc, where c≥1, according to the order of their acquisition time from the current time to the earliest. SS12: Determine whether environmental parameter B1 is an interfering parameter of analysis data A1. The determination criteria are as follows: A discrete point filtering algorithm is used to process the monitoring values C1, C2, ..., Cc, and all remaining monitoring values after data processing are mapped onto a line graph. In the line graph, the horizontal axis represents the acquisition time, and the vertical axis represents the monitoring value. All remaining monitoring values after data processing are mapped onto the line graph in order from the acquisition time to the current time, from the furthest to the nearest. If all monitored values on the line graph show a decreasing or increasing trend from left to right, then environmental parameter B1 is determined to be an interference parameter of analytical data A1; otherwise, environmental parameter B1 is determined not to be an interference parameter of analytical data A1. The judgment is explained as follows: Analysis data A1 was generated under the preset collection conditions of a certain voice command. Among the preset collection conditions, it is stated that the confidence level needs to show a continuous downward trend on the line graph. The continuous downward trend satisfies monotonicity and is continuous. If the overall monitoring value of a certain environmental parameter also satisfies monotonicity and is continuous, then whether it is continuously rising or continuously falling, it means that it is synchronized with the trend of confidence level change. SS13: Determine whether environmental parameters B2, B3, ..., Bb are interference parameters of analysis data A1 in SS11 to SS12. After the determination, re-mark all interference parameters that are determined to be interference parameters of analysis data A1 as D1, D2, ..., Dd, where 1≤d≤b. The subscript of the marking indicates the order in which they are determined to be interference parameters. The smaller the subscript, the earlier it is determined to be an interference parameter. S13: Select several voice command issuance times from the third data within the analysis data A1 as gradient inflection points of analysis data A1. Combine these with the acquisition times corresponding to the farthest and closest monitoring values within analysis data A1 to determine the interference intervals G1, G2, ..., Gg, where 1 ≤ g ≤ f + 1, and f is the total number of selected gradient inflection points in analysis data A1. The specific details are as follows: SS21: Extract the confidence scores of all voice commands from the third data in the analysis data A1, and label all extracted confidence scores as E1, E2, ..., Ee in order of their distance from the current time, from farthest to closest, where e≥1; SS22: Compare the confidence scores E1 and E2. If the confidence score E1 is greater than E2 and the value of confidence score E1-E2 is greater than or equal to the preset standard gradient difference, then the moment when the voice command is issued corresponding to confidence score E2 is selected as a gradient inflection point. Otherwise, no processing is performed. The value of the standard gradient difference is represented by a relatively obvious downward trend. SS23: Following the steps in SS22, compare the confidence levels E2 and E3, E3 and E4, ..., Ee-1 and Ee in sequence. After the comparison, relabel all gradient inflection points in the order they were selected, from first to last, as F1, F2, ..., Ff, where 1≤f <e; SS24: Traverse the second data in the analysis data A1, mark the collection time corresponding to the monitoring value farthest from the current time as the starting point, and mark the collection time corresponding to the monitoring value closest to the current time as the ending point. Combine the gradient inflection points F1, F2, ..., Ff to divide into several interference intervals, and mark them in chronological order from first to last as G1, G2, ..., Gg, 1≤g≤f+1. Among them, interference interval G1 corresponds to the time from the starting point to the gradient inflection point F1, G2 corresponds to the time from the next time of gradient inflection point F1 to gradient inflection points F2, G3, ..., Gg, and so on. Additional explanation: The starting point and gradient inflection point F1 are called the starting point and ending point of the interference interval G1, respectively, and the interference intervals G2, G3, ..., Gg are called in the same way. S14: Calculate the average interference weights of the interference parameters D1, D2, ..., Dd within the analysis data A1. The calculation is as follows: SS31: Extract the maximum and minimum monitoring values from all monitoring values of interference parameter D1 in the second data, calculate the difference between the two extracted monitoring values, and take the difference as the parameter variation amplitude H1-1 of interference parameter D1 in interference interval G1. Similarly, calculate the parameter variation amplitude H1-2, H1-3, ..., H1-g of interference parameter D1 in interference intervals G2, G3, ..., Gg. SS32: Obtain the confidence scores of all voice commands issued within the interference interval G1 from the third data, extract the highest and lowest confidence scores, and calculate the difference to obtain the confidence score variation amplitude I1 within the interference interval G1. SS33: Utilizing Formulas The interference ratio J1 of the interference parameter D1 in the interference interval G1 is calculated and obtained. The interference ratio J1 is used to characterize the influence ratio of the confidence change amplitude I1 caused by the interference parameter D1 under the combined action of several interference parameters. In the formula, Hi-1 represents each of the parameter change amplitudes H1-1, H1-2, ..., H1-g. Taking the square root of the square of each parameter change amplitude in the formula is used to eliminate the calculation deviation caused by the continuous increase or decrease of the monitored value. SS34: Calculate and obtain the interference percentages J2, J3, ..., Jg of interference parameter D1 in the interference intervals G2, G3, ..., Gg in sequence according to SS33; SS35: A discrete point filtering algorithm is used to process J1, J2×η1, J3×η2, ..., Jg×ηg-1 to obtain the average value of all remaining data after data processing. The average value is then calibrated as the average interference weight of interference parameter D1, where η1, η2, ..., ηg-1 are preset compensation coefficients for interference parameter D1 as the value of the marked index of the interference interval increases. SS36: Calculate and obtain the average interference weights of interference parameters D2, D3, ..., Dd in sequence according to SS31 to SS35; S15: Calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data point in the analysis data A1. The calculations are as follows: SS41: Constructing an equation based on the interference interval G1 In the equation, a1, a2, and a3 are the interference parameter coefficients, confidence level variation coefficients, and interval duration coefficients to be solved, respectively; L1 is the interval duration of the interference interval G1; Hj-1 refers to each of the parameter variation amplitudes H1-1, H2-1, ..., Hd-1; H2-1, ..., Hd-1 are the parameter variation amplitudes of the interference parameters D2, D3, ..., Dd in the interference interval G1, calculated when calculating the average interference weights of the interference parameters D2, D3, ..., Dd. In addition, in the above-constructed equation, the left side of the equation represents the comprehensive interference risk value that induces speech recognition anomalies within the interference interval G1; the right side of the equation represents the system's carrying capacity for the corresponding comprehensive interference risk value. SS42: Construct equations sequentially based on interference intervals G2, G3, ..., Gf+1, following the pattern in SS41. The equation constructed based on interference interval G2 is... In the formula, β1, β2, and β3 are the preset first, second, and third compensation coefficients, respectively. Their values are set based on the interval duration of interference interval G2 divided by the sum of the interval durations of interference intervals G1, G2, ..., Gf+1. Equations are constructed based on interference intervals G3, G4, ..., Gf+1, and so on, to assign weights to interference intervals of different durations, so as to avoid excessive influence of intervals with too short durations on the coefficient solution. SS43: By constructing a system of three linear equations in three variables based on any three equations of the interference intervals G1, G2, ..., Gg, we can obtain C. g 3 Solve each of the three systems of linear equations in three variables to obtain C. g 3 The values of a1, a2 and a3; SS44: Using a discrete point filtering algorithm for C f+1 3The values of a1 are processed to obtain the average value of all remaining a1 values after data processing. The average value is then calibrated as the interference coefficient of the first data in the analysis data A1. SS45: Apply discrete point filtering algorithms to C according to SS43 to SS44 respectively. f+1 3 The values of a2 and a3 are processed to obtain the average value of all remaining a2 and a3 values after data processing. The average value is then labeled as the confidence coefficient and duration coefficient of the first data in the analysis data A1. S16: Calculate the interference weights of all interference parameters in the analysis data A2, A3, ..., Aa according to S14; calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data in the analysis data A2, A3, ..., Aa according to S15. S17: Based on the first data in the analysis data A1, A2, ..., Aa, construct several preloaded sets, labeled M1, M2, ..., Mm, where 1 ≤ m ≤ a, with the following contents: Extract the first data from each of the analysis data A1, A2, ..., Aa and remove duplicates. Add all the processing algorithms contained in each of the remaining first data after deduplication to an empty set to obtain the corresponding preloaded set. S18: Generate matching data for the preloaded sets M1, M2, ..., Mm according to the preset generation steps. The generated content is as follows: SS51: Extract all analysis data from analysis data A1, A2, ..., Aa that match the preloaded set M1, and relabel them as N1, N2, ..., Nn, where 1 ≤ n <a; SS52: Obtain all interference parameters of the analysis data N1, add all the interference parameters to an empty set to obtain the interference set of the analysis data N1, and similarly obtain the interference sets of the analysis data N2, N3, ..., Nn respectively. Iterate through all the obtained interference sets and select the interference set with the most consistent values as the associated set of the preloaded set M1. SS53: Relabel all analytical data within the analytical data N1, N2, ..., Nn, whose interference set is the association set, as Q1, Q2, ..., Qq, 1≤q≤n; SS54: For each interference parameter in the association set, obtain the average interference weight of the interference parameter in the analysis data Q1, Q2, ..., Qq respectively, and calculate its mean using the summation and averaging formula. Use the mean as the standard interference weight of the interference parameter in the association set of the preloaded set M1. SS55: Calculate the mean value of the interference coefficient of the first data in the analysis data Q1, Q2, ..., Qq using the summation and averaging formula, and use the mean value as the mean interference coefficient of the preloaded set M1. Similarly, calculate the mean value of the confidence coefficient and the mean value of the duration coefficient of the first data in the analysis data Q1, Q2, ..., Qq, and use the mean value as the mean confidence coefficient and the mean value of the duration coefficient of the preloaded set M1, respectively. SS56: Generate matching data for preloaded set M1 based on the average interference coefficient, average confidence coefficient, average duration coefficient, and the standard interference weights of all interference parameters in the associated set of preloaded set M1. SS57: Calculate and obtain the matching data of the preloaded sets M2, M3, ..., Mm in sequence according to SS51 to SS56; The data analysis module transmits the pre-loaded matching data of sets M1, M2, ..., Mm to the voice scheduling module for storage.
[0018] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.
[0019] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction, characterized in that, include: The environmental monitoring module is used to collect environmental monitoring data of the target cabin in real time. The voice dispatch module is used to determine whether the voice command is a local or cloud-based recognition subject based on the network connectivity status of the target cabin when the voice command is collected, according to the voice command issued by the personnel stationed in the target cabin. The network connectivity status is either online or offline. If the target cabin's network connectivity is in a connected state, a preliminary determination is made of the entity that recognizes and executes the voice command. Based on the preliminary determination result, it is determined whether the entity that recognizes and executes the voice command is a local entity. If it is not a local entity, an advanced determination is made of the voice command. Based on the advanced determination result, it is determined whether the entity that recognizes and executes the voice command is a local entity or a cloud entity. If the network connectivity status of the target cabin is no network, then the recognition of the voice command is directly determined to be local recognition. The data analysis module is used to collect and generate a fixed amount of analytical data that can be analyzed. The content of any analytical data collected and generated is as follows: when any voice command is detected, if the voice command meets the preset collection conditions, the analytical data of the voice command is generated. The analytical data includes the first data of the voice command obtained from the cloud recognition module, the second data of the voice command obtained from the environmental monitoring module, and the third data of the voice command obtained from the voice scheduling module. The data analysis module is also used to analyze all the generated analysis data after the amount of collected and generated analysis data reaches a preset fixed amount, and generate matching data for several preloaded sets. The cloud recognition module is used to input each voice command received by the cloud recognition execution entity into the pre-trained cloud big model for recognition and parsing. The cloud big model outputs the corresponding voice text. After the recognition is completed, it automatically records and stores all the processing algorithms used by the cloud big model to recognize the voice command. The local recognition module is used to input each received voice command, which is performed by a local recognition entity, into a pre-trained local large model to recognize and output the corresponding speech text.
2. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 1, characterized in that, The initial determination of the entity responsible for recognizing the voice command is as follows: The PESQ method is used to output the confidence level of the voice command. Starting from the time the voice command is issued, a preset comparison duration P4 is traced backward to define the comparison period. All voice commands issued by personnel within the target cabin at times within this comparison period are acquired. The confidence level difference between the two voice commands with the closest issuance times is calculated sequentially. If any calculated confidence level difference is greater than or equal to a preset confidence level difference threshold, an advanced determination is performed to determine the entity responsible for recognizing the voice command. If all calculated confidence level differences are less than the preset confidence level difference threshold, the entity responsible for recognizing the voice command is determined to be a local entity.
3. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 1, characterized in that, The preset data collection conditions are as follows: Starting from the moment the voice command is issued, a preset environmental change observation duration P1 is traced back to determine the corresponding end time, thereby defining the collection period. All voice commands issued within the collection period are acquired. If the recognition execution entity of all voice commands is local recognition, and the confidence of all voice commands shows a continuous downward trend on the line graph, and the voice command is a correction of the previous voice command, then the voice command is determined to meet the preset collection conditions. Here, the voice command is a correction of the previous voice command, which is a subsequent voice command issued by the stationed personnel to correct the recognition abnormality caused by the inaccurate recognition result of the previous voice command.
4. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 1, characterized in that, The content of the matching data generated for several preloaded collections is as follows: S11: Label all generated analysis data as A1, A2, ..., Aa, where a≥1; label all environmental parameters selected by the management personnel as B1, B2, ..., Bb, where b≥1. S12: According to the preset judgment steps, determine whether the environmental parameters B1, B2, ..., Bb are interference parameters of the analysis data A1, and re-mark all interference parameters that are judged to be analysis data A1 as D1, D2, ..., Dd, 1≤d≤b; S13: Select several voice command issuance times from the third data in the analysis data A1 as gradient inflection points of the analysis data A1. Combine the acquisition times corresponding to the monitoring values that are farthest and closest to the current time in the analysis data A1 to determine the interference intervals G1, G2, ..., Gg, where 1≤g≤f+1, and f is the total number of gradient inflection points of the selected analysis data A1. S14: Calculate the average interference weights of the interference parameters D1, D2, ..., Dd within the analysis data A1; S15: Calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data in the analysis data A1; S16: Calculate the interference weights of all interference parameters in the analysis data A2, A3, ..., Aa according to S14; calculate the interference coefficient, confidence coefficient, and duration coefficient of the first data in the analysis data A2, A3, ..., Aa according to S15. S17: Based on the first data in the analysis data A1, A2, ..., Aa, construct several preloaded sets, labeled as M1, M2, ..., Mm, where 1≤m≤a; S18: Generate matching data for preloaded sets M1, M2, ..., Mm according to preset generation steps, and transfer the matching data of preloaded sets M1, M2, ..., Mm to the voice scheduling module for storage. The generation steps are as follows: SS51: Extract all analysis data from analysis data A1, A2, ..., Aa that match the preloaded set M1, and relabel them as N1, N2, ..., Nn, where 1 ≤ n <a; SS52: Obtain all interference parameters of the analysis data N1, add all the interference parameters to an empty set to obtain the interference set of the analysis data N1, and similarly obtain the interference sets of the analysis data N2, N3, ..., Nn respectively. Iterate through all the obtained interference sets and select the interference set with the most consistent values as the associated set of the preloaded set M1. SS53: Relabel all analytical data within the analytical data N1, N2, ..., Nn, whose interference set is the association set, as Q1, Q2, ..., Qq, 1≤q≤n; SS54: For each interference parameter in the association set, obtain the average interference weight of the interference parameter in the analysis data Q1, Q2, ..., Qq respectively, and calculate its mean using the summation and averaging formula. Use the mean as the standard interference weight of the interference parameter in the association set of the preloaded set M1. SS55: Calculate the mean value of the interference coefficient of the first data in the analysis data Q1, Q2, ..., Qq using the summation and averaging formula, and use the mean value as the mean interference coefficient of the preloaded set M1. Similarly, calculate the mean value of the confidence coefficient and the mean value of the duration coefficient of the first data in the analysis data Q1, Q2, ..., Qq, and use the mean value as the mean confidence coefficient and the mean value of the duration coefficient of the preloaded set M1, respectively. SS56: Generate matching data for preloaded set M1 based on the average interference coefficient, average confidence coefficient, average duration coefficient, and the standard interference weights of all interference parameters in the associated set of preloaded set M1. SS57: Calculate and obtain matching data for the preloaded sets M2, M3, ..., Mm in sequence according to SS51 to SS56.
5. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 4, characterized in that, S12, the determination steps are as follows: SS11: Extract all monitoring values of environmental parameter B1 from the second data set of analysis data A1, and label all extracted monitoring values as C1, C2, ..., Cc, where c≥1, according to the order of their acquisition time from the current time to the earliest. SS12: Determine whether environmental parameter B1 is an interference parameter of analysis data A1. The determination is as follows: Use a discrete point filtering algorithm to process the monitoring values C1, C2, ..., Cc, and map all the remaining monitoring values after data processing onto a line graph. If all monitored values on the line graph show a decreasing or increasing trend from left to right, then environmental parameter B1 is determined to be an interference parameter of analytical data A1; otherwise, environmental parameter B1 is determined not to be an interference parameter of analytical data A1. SS13: Determine whether environmental parameters B2, B3, ..., Bb are interference parameters of analysis data A1 in SS11 to SS12. After the determination, re-label all interference parameters determined to be analysis data A1 as D1, D2, ..., Dd, where 1≤d≤b.
6. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 4, characterized in that, S13, the details are as follows: SS21: Extract the confidence scores of all voice commands from the third data in the analysis data A1, and label all extracted confidence scores as E1, E2, ..., Ee in order of their distance from the current time, e≥1; SS22: Compare the confidence scores E1 and E2. If the confidence score E1 is greater than E2 and the value of confidence score E1-E2 is greater than or equal to the preset standard gradient difference, then select the time of speech command issuance corresponding to confidence score E2 as a gradient inflection point; otherwise, no processing is performed. SS23: Following the steps in SS22, compare the confidence levels E2 and E3, E3 and E4, ..., Ee-1 and Ee in sequence. After the comparison, relabel all gradient inflection points in the order they were selected, from first to last, as F1, F2, ..., Ff, where 1≤f <e; SS24: Traverse the second data in the analysis data A1, mark the collection time corresponding to the monitoring value farthest from the current time as the starting point, and mark the collection time corresponding to the monitoring value closest to the current time as the ending point. Combine the gradient inflection points F1, F2, ..., Ff to divide into several interference intervals, and mark them in chronological order from first to last as G1, G2, ..., Gg, 1≤g≤f+1. Among them, interference interval G1 corresponds to the time from the starting point to the gradient inflection point F1, G2 corresponds to the time from the next time of gradient inflection point F1 to gradient inflection points F2, G3, ..., Gg, and so on.
7. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 4, characterized in that, S14, the steps for calculating the average weight of the interference parameters D1, D2, ..., Dd within the analysis data A1 are as follows: SS31: Extract the maximum and minimum monitoring values from all monitoring values of interference parameter D1 in the second data, calculate the difference between the two extracted monitoring values, and take the difference as the parameter variation amplitude H1-1 of interference parameter D1 in interference interval G1. Similarly, calculate the parameter variation amplitude H1-2, H1-3, ..., H1-g of interference parameter D1 in interference intervals G2, G3, ..., Gg. SS32: Obtain the confidence scores of all voice commands issued within the interference interval G1 from the third data, extract the highest and lowest confidence scores, and calculate the difference to obtain the confidence score variation amplitude I1 within the interference interval G1. SS33: Utilizing Formulas Calculate the interference percentage J1 of interference parameter D1 in interference interval G1, where Hi-1 represents each of the parameter variation amplitudes H1-1, H1-2, ..., H1-g; SS34: Calculate and obtain the interference percentages J2, J3, ..., Jg of interference parameter D1 in the interference intervals G2, G3, ..., Gg in sequence according to SS33; SS35: A discrete point filtering algorithm is used to process J1, J2×η1, J3×η2, ..., Jg×ηg-1 to obtain the average value of all remaining data after data processing. The average value is then calibrated as the average interference weight of interference parameter D1, where η1, η2, ..., ηg-1 are preset compensation coefficients for interference parameter D1 as the value of the marked index of the interference interval increases. SS36: Calculate and obtain the average interference weights of interference parameters D2, D3, ..., Dd in sequence according to SS31 to SS35.
8. The multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction according to claim 7, characterized in that, S15, the calculation content is as follows: SS41: Constructing an equation based on the interference interval G1 In the equation, a1, a2, and a3 are the interference parameter coefficients, confidence level variation coefficients, and interval duration coefficients to be solved, respectively; L1 is the interval duration of the interference interval G1; Hj-1 refers to each of the parameter variation amplitudes H1-1, H2-1, ..., Hd-1; H2-1, ..., Hd-1 are the parameter variation amplitudes of the interference parameters D2, D3, ..., Dd in the interference interval G1, calculated when calculating the average interference weights of the interference parameters D2, D3, ..., Dd. SS42: Construct equations sequentially based on interference intervals G2, G3, ..., Gf+1, following the pattern in SS41. The equation constructed based on interference interval G2 is... In the formula, β1, β2, and β3 are the preset first, second, and third compensation coefficients, respectively; SS43: By constructing a system of three linear equations in three variables based on any three equations of the interference intervals G1, G2, ..., Gg, we can obtain C. g 3 Solve each of the three systems of linear equations in three variables to obtain C. g 3 The values of a1, a2 and a3; SS44: Using a discrete point filtering algorithm for C f+1 3 The values of a1 are processed to obtain the average value of all remaining a1 values after data processing. The average value is then calibrated as the interference coefficient of the first data in the analysis data A1. SS45: Apply discrete point filtering algorithms to C according to SS43 to SS44 respectively. f+1 3 The values of a2 and a3 are processed to obtain the average value of all remaining a2 and a3 values after data processing. The average value is then labeled as the confidence coefficient and duration coefficient of the first data in the analysis data A1.
9. A multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 4, characterized in that, The advanced determination of the voice command is as follows: Starting from the moment the voice command is issued, a preset observation duration P3 is traced back to define the observation period. Based on the observation period, all interference parameters relative to the voice command are selected from all environmental parameters selected by the management personnel. All interference parameters within the observation period are integrated into an interference set. The interference set is used as a search condition. If an associated set that matches the interference set is found, the matching data of the preloaded set corresponding to the associated set is obtained. Otherwise, it is determined that the recognition execution subject of the voice command is local recognition. The average interference coefficient of each interference parameter is obtained from the matching data, and then the formula is used. Calculate and obtain the predicted recognition duration T1 based on the voice command, where Uu refers to each interference parameter in the observation period, λu represents the interference standard weight of the corresponding interference parameter from the matching data, X1, X2, and X3 are the average interference coefficient, average confidence coefficient, and average duration coefficient contained in the matching data, respectively, and v is the total number of interference parameters determined to be interference parameters in the observation period. If T1≈0, the entity executing the voice command recognition is determined to be cloud-based recognition; otherwise, the entity executing the voice command recognition is determined to be local recognition. At this time, a task timing instruction is generated based on the calculated predicted recognition duration T1, and the task timer is started. The task timer is initially assigned a value of 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1, the entity executing the recognition of all voice commands collected subsequently is automatically determined to be cloud-based recognition, and no network status or any subsequent determinations are made. The pre-fetching recognition time T1 and the pre-loaded set are simultaneously transmitted to the cloud recognition module.
10. A multi-parameter voice control system for a high-altitude intelligent cabin integrating voice interaction as described in claim 9, characterized in that, After receiving the pre-fetch recognition duration T1 and the preloaded set, the cloud recognition module synchronously starts the task timer. The task timer is initially set to 0 and executes a counting logic that increments by 1 each time. When the cumulative timer duration reaches T1×ε1, the preloading operation is automatically executed to preload all the processing algorithms contained in the preloaded set into the runtime cache.