Instruction conversion control method and system based on voice recognition
By building an analysis model for the degree of operation threat and abnormal state of content parsing, combined with an operation intention risk assessment model, the problem of failing to assess the threat of command execution in voice command control is solved, thereby improving the security and stability of the system.
Patent Information
- Application Number
- CN202510891419.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing technologies lack in-depth analysis of voice commands in voice command control and fail to effectively assess the threats and conflicts posed by command execution to current operations, resulting in unstable system operation and high risk of failure.
By building an operation threat level analysis model and a content parsing abnormal state analysis model, combined with an operation intention risk assessment model, a comprehensive analysis and evaluation of the voice commands to be executed is conducted, and a decision is made on whether to convert the commands based on the evaluation results.
It improves the security and stability of system operation, reduces the risk of failure during instruction conversion, and realizes intelligent decision-making and safety prediction.
Smart Images

Figure CN120612927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of command conversion control, and in particular to a command conversion control method and system based on speech recognition. Background Art
[0002] In the process of intelligent system interaction and automated control, command conversion control driven by voice recognition technology is a critical step in achieving efficient human-machine collaboration and intelligent device response. As the core vehicle for human-machine interaction, the accuracy of voice command parsing, operational safety, and execution reliability directly impact the safety and stability of intelligent system operations. For example, in intelligent production line control, issues with command parsing accuracy can trigger the execution of illegal commands, or lead to production line interruptions due to a failure to assess the threat posed by command execution to the current production process. In intelligent traffic control, scheduling errors can also occur due to incomplete command parsing or failure to consider execution threats. Traditional voice command control methods focus primarily on command recognition accuracy, simply converting recognition results directly into control operations. They lack in-depth analysis of voice commands and a comprehensive assessment of their impact on current operations. Consequently, traditional voice command control methods struggle to address the risk management requirements of command execution in complex scenarios and cannot guarantee reliable system operation within human-machine interaction and automated processes.
[0003] At the same time, the existing technology only relies on text accuracy matching of voice command content to decide whether to execute control commands, ignoring the quantitative analysis of the semantic integrity and compliance of voice commands, and does not take into account the various conflicting factors between the commands to be executed and the current running operations that may cause equipment operation to be interrupted; and the existing technology does not incorporate command parsing anomalies and operation threats into the risk judgment of operational intentions during command conversion, making the operational risk identification during command conversion control too one-sided and the command execution too blind.
[0004] In order to solve these problems, the present application designs a command conversion control method and system based on speech recognition. Summary of the Invention
[0005] To overcome the shortcomings and deficiencies of existing technologies, the present invention provides a voice recognition-based command conversion control method and system. This method comprehensively analyzes the operational threat level and content parsing anomalies posed by the pending voice command to the currently running operation; constructs an operation intention risk assessment model to assess the operation intention risk of the pending voice command; and converts the pending voice command based on the operation intention risk assessment results. This improves the security and stability of system operation and reduces the risk of failure during the command conversion process.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, an embodiment of the present invention provides a method for controlling command conversion based on speech recognition, comprising the following steps: S1. Obtaining voice command data to be executed, historical accident data, and current operation data; S2. Importing the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; S3. Importing the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; S4. Build an operation intention risk assessment model, import the analysis results of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis results of the abnormal state of the content analysis into the operation intention risk assessment model, and evaluate the operation intention risk of the voice command to be executed; S5. Convert the voice command to be executed based on the risk assessment result of the operation intention.
[0007] In the preferred technical solution of the present invention, step S2 analyzes the operational threat level caused by the voice command to be executed to the currently running operation, including the following specific steps: S21, extracting the voice command data to be executed and the current operation data; S22. Based on the pending voice command data and the current running operation data, construct an operation threat level analysis model to analyze the operational threat level posed by the pending voice command to the currently running operation, and obtain an analysis result of the operational threat level posed by the pending voice command to the currently running operation; The calculation formula for the operation threat level is: ; Where Cz is the operational threat level caused by the pending voice command to the currently running operation, ct is the analysis result of the compatibility conflict threat level caused by the pending voice command to the currently running operation, and yc is the analysis result of the abnormal urgency threat level caused by the pending voice command to the currently running operation.
[0008] In the preferred technical solution of the present invention, the process of constructing the running operation threat level analysis model in step S22 includes the following specific steps: S221. Analyzing the compatibility conflict threat level caused by the voice command to be executed on the currently running operation based on the voice command data to be executed and the currently running operation data, thereby obtaining a compatibility conflict threat level analysis result caused by the voice command to be executed on the currently running operation; S222. Based on the voice instruction data to be executed and the current operation data, analyze the abnormal urgency threat degree caused by the voice instruction to be executed to the currently running operation, and obtain an analysis result of the abnormal urgency threat degree caused by the voice instruction to be executed to the currently running operation.
[0009] In the preferred technical solution of the present invention, step S3 analyzes the abnormal state of content parsing of the voice instruction to be executed, including the following specific steps: S31, extracting voice command data to be executed and historical accident data; S32. Based on the voice command data to be executed and the historical accident data, a command content parsing abnormal state analysis model is constructed, and the abnormal state of the content parsing of the voice command to be executed is analyzed to obtain an analysis result of the abnormal state of the content parsing of the voice command to be executed; The calculation formula for content parsing abnormal status is: ; Where Nr is the abnormal state of content analysis of the voice command to be executed, wz is the abnormal state analysis result of the semantic integrity of the voice command to be executed, and hg is the abnormal state analysis result of the semantic compliance of the voice command to be executed.
[0010] In the preferred technical solution of the present invention, the process of constructing the instruction content parsing abnormal state analysis model in step S32 includes the following specific steps: S321. Analyze the semantic integrity abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain an analysis result of the semantic integrity abnormality state of the voice command to be executed; S322: Analyze the semantic compliance abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain an analysis result of the semantic compliance abnormality state of the voice command to be executed.
[0011] In the preferred technical solution of the present invention, the operation intention risk assessment model is constructed in step S4, including the following specific steps: S41, extracting and analyzing the threat level of the voice command to be executed to the currently running operation, and the abnormal state analysis result of the content parsing of the voice command to be executed; S42. Assess the risk of the operation intention of the voice command to be executed based on the analysis result of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis result of the abnormal state of the content analysis; The evaluation formula for operation intention risk is: ; Where ZY is the operation intention risk of the voice command to be executed, Cz is the analysis result of the operation threat degree caused by the voice command to be executed to the current operation, and Nr is the analysis result of the abnormal state of the content analysis of the voice command to be executed. They are the operation threat impact weight and content analysis anomaly impact weight.
[0012] In the preferred technical solution of the present invention, in step S5, the voice command to be executed is converted according to the operation intention risk assessment result, including the following specific steps: S51. Obtaining a risk assessment result of the operation intention of the voice command to be executed; S52. Preset an operation intention risk threshold. When the operation intention risk assessment result of the voice instruction to be executed is greater than or equal to the operation intention risk threshold, refuse to convert the voice instruction to be executed; when the operation intention risk assessment result of the voice instruction to be executed is less than the operation intention risk threshold, convert the voice instruction to be executed.
[0013] In a second aspect, an embodiment of the present invention further provides a command conversion control system based on speech recognition, including: Data acquisition module, used to obtain voice command data to be executed, historical accident data and current operation data; An operation threat level analysis module is used to import the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; The command content parsing abnormal state analysis module is used to import the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; The operation intention risk assessment module is used to build an operation intention risk assessment model. The operation threat analysis results of the voice command to be executed on the currently running operation and the abnormal state analysis results of the content analysis are imported into the operation intention risk assessment model to assess the operation intention risk of the voice command to be executed; The command conversion control module is used to convert the voice commands to be executed based on the risk assessment results of the operation intention; The control module is used to control the operation of the data acquisition module, the operation threat level analysis module, the instruction content parsing abnormal state analysis module, the operation intention risk assessment module and the instruction conversion control module.
[0014] Compared with the prior art, the present invention has the following advantages and beneficial effects: The present invention analyzes the operational threat level posed by a pending voice command to the currently running operation; analyzes the abnormal state of the content parsing of the pending voice command; constructs an operational intention risk assessment model, imports the operational threat level analysis results and the abnormal state analysis results of the content parsing of the pending voice command to the currently running operation into the operational intention risk assessment model, and then assesses the operational intention risk of the pending voice command; based on the operational intention risk assessment results, the pending voice command is converted. This enables safe prejudgment and intelligent decision-making for command conversion control, improves the safety and stability of system operation, and reduces the risk of failure during the command conversion process. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 Schematic diagram of the overall flow of the command conversion control method based on speech recognition of the present invention; Figure 2 This is a workflow diagram of step S2 in the command conversion control method based on speech recognition of the present invention; Figure 3 This is a workflow diagram of step S3 in the command conversion control method based on speech recognition of the present invention; Figure 4 It is a structural diagram of the command conversion control system based on speech recognition of the present invention. DETAILED DESCRIPTION
[0016] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0017] Example 1
[0018] like Figure 1 As shown, this embodiment provides a command conversion control method based on voice recognition, which specifically includes the following steps: S1. Obtaining voice command data to be executed, historical accident data, and current operation data; S2. Importing the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; S3. Importing the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; S4. Build an operation intention risk assessment model, import the analysis results of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis results of the abnormal state of the content analysis into the operation intention risk assessment model, and evaluate the operation intention risk of the voice command to be executed; S5. Convert the voice command to be executed based on the risk assessment result of the operation intention.
[0019] In this embodiment, if Figure 2 As shown, in step S2, the operational threat level caused by the voice command to be executed to the currently running operation is analyzed, including the following specific steps: S21, extracting the voice command data to be executed and the current operation data; S22. Based on the pending voice command data and the current running operation data, construct an operation threat level analysis model to analyze the operational threat level posed by the pending voice command to the currently running operation, and obtain an analysis result of the operational threat level posed by the pending voice command to the currently running operation; The calculation formula for the operation threat level is: ; Where Cz is the operational threat level caused by the pending voice command to the currently running operation, ct is the analysis result of the compatibility conflict threat level caused by the pending voice command to the currently running operation, and yc is the analysis result of the abnormal urgency threat level caused by the pending voice command to the currently running operation.
[0020] For example, in this embodiment, the sum of the compatibility conflict threat level analysis results and the abnormal urgency threat level analysis results is used to reflect the total contribution of the two types of threats to the operational threat level caused by the voice command to be executed to the currently running operation, thereby reflecting the basic superposition of the operational threat level; the denominator middle, It is used to describe the difference between the compatibility conflict threat level and the abnormal urgency threat level. The greater the difference, the larger the denominator value, which can suppress the unreasonable amplification of the operational threat level caused by the extreme difference between the two types of threats. By adding 2, the denominator is avoided from being zero. The value range of Cz is constrained to be within the interval of -1 to 1, and after being mapped by the exponential function, it not only retains the superposition trend of the two types of threats, but also avoids the occurrence of a single threat that over-dominates the final operational threat level due to simple linear operations. In addition, by subtracting 1, the baseline alignment of the operational threat level is ensured, that is, when there is no compatible conflict threat and abnormal urgency threat, the value of Cz is 0. Specifically, this embodiment comprehensively calculates the compatible conflict threat and the abnormal urgency threat, which not only ensures the superposition effect of the two types of threats, but also uses the denominator to This dynamically constrains the extreme interference caused by the difference between the two types of threats. By introducing an exponential function, the calculation formula provided in this embodiment can meet the actual evolutionary law of threat level from zero to a gradual accumulation, and can accurately reflect the comprehensive operational threat level dominated by the dual dimensions of compatible conflict threats and abnormal urgency threats.
[0021] In this embodiment, the process of constructing the operational threat level analysis model in step S22 includes the following specific steps: S221. Analyzing the compatibility conflict threat level caused by the voice command to be executed on the currently running operation based on the voice command data to be executed and the currently running operation data, thereby obtaining a compatibility conflict threat level analysis result caused by the voice command to be executed on the currently running operation; The calculation formula for the compatibility conflict threat level is: ; Where ct is the compatibility conflict threat degree caused by the voice command to be executed to the currently running operation, vi is the operation action feature vector sequence parsed from the voice command to be executed in the voice command data to be executed, and vo is the operation action feature vector sequence of the current running operation in the current running operation data. is the squared Euclidean distance between vi and vo, is the covariance of vi and vo, wherein the calculation formula of the squared Euclidean distance and the covariance calculation formula belong to the prior art and will not be repeated here; For example, this embodiment analyzes the compatibility conflict threat level based on the difference and synergy of operation actions. The core is the sequence of the operation action feature vectors of the voice command to be executed and the current operation. The difference in the operation actions of the two is measured by calculating the square Euclidean distance between the two. The greater the difference, the higher the compatibility conflict potential. It reflects the synergy of the operational action feature vector sequences of the two. When the synergy is low, the compatibility conflict threat will be amplified. Furthermore, this embodiment uses an exponential function and an inverse operation to perform a linear operation on the action difference and synergy between the voice instruction to be executed and the current operation, and then converts the linear operation result into a value with a specific distribution law through an exponential function, thereby compressing the influence of the action difference and synergy to a reasonable range, so that the final compatibility conflict threat level can smoothly and reasonably reflect the conflict relationship between the two. At the same time, it enhances the sensitivity of the compatible conflict threat level calculation formula provided by this embodiment to the changes in difference and synergy, and highlights the discrimination of the conflict degree under different action feature vector sequences. Specifically, in this embodiment, the acquisition of the operational action feature vector sequence of the voice instruction to be executed and the current operation is to perform semantic analysis on the voice instruction, extract the key elements of the operational action, such as action type, execution order, action object, etc., and convert it into a feature vector containing action type, action object, and parameters through encoding. The operational action feature vector sequence of the current operation can also be obtained by extracting the same type of feature encoding of the corresponding operational action from the real-time monitoring data of the current operation.
[0022] S222: Analyze the degree of abnormal urgency threat posed by the voice command to be executed to the currently running operation based on the pending voice command data and the currently running operation data, and obtain an analysis result of the degree of abnormal urgency threat posed by the pending voice command to the currently running operation; The calculation formula for abnormal urgency threat level is: ; Where yc is the degree of abnormal urgency threat posed by the pending voice command to the currently running operation, Ci is the number of operation steps parsed from the pending voice command in the pending voice command data, Ti is the required execution time limit of the pending voice command in the pending voice command data, To is the remaining execution time of the current running operation in the current running operation data, and Ro is the system resource usage rate of the current running operation in the current running operation data; For example, this embodiment analyzes the abnormal urgency threat level based on the urgency of the voice command for time and resources. Analyze the impact of the pending voice command execution on the time resources of the current operation. The more steps in the operation and the larger the time difference, the more significant the impact on the time resources of the current operation. It is used to reflect the resource adaptability of the current operation within the required execution time limit of the voice instruction to be executed. The lower the system resource utilization rate of the current operation, the stronger its resource adaptability within the required execution time limit of the voice instruction to be executed, which can alleviate the urgency of resource utilization. Specifically, this embodiment also uses an exponential function to perform a nonlinear transformation on the calculation results of factors such as the number of operation steps and time difference. The time resource impact reflected by the number of operation steps, the required execution time limit and the remaining execution time is converted into a value that conforms to a specific attenuation law after being processed by an exponential function, so that the calculation result of the abnormal urgency threat level can be kept within a reasonable range, and can highlight the characteristics that the abnormal urgency threat rises rapidly and then tends to be stable when the time difference is large, which is more in line with the evolution law of abnormal urgency threats in the actual instruction conversion control process. Furthermore, in this embodiment, the number of operation steps is obtained by parsing the voice instruction semantics and disassembling the execution process according to the instruction semantics to obtain the number of operation steps; the required execution time limit is obtained by extracting the time requirement expression from the voice instruction and converting it into a numerical value.
[0023] In this embodiment, if Figure 3 As shown, in step S3, the abnormal state of content parsing of the voice command to be executed is analyzed, including the following specific steps: S31, extracting voice command data to be executed and historical accident data; S32. Based on the voice command data to be executed and the historical accident data, a command content parsing abnormal state analysis model is constructed, and the abnormal state of the content parsing of the voice command to be executed is analyzed to obtain an analysis result of the abnormal state of the content parsing of the voice command to be executed; The calculation formula for content parsing abnormal status is: ; Where Nr is the abnormal state of content analysis of the voice command to be executed, wz is the abnormal state analysis result of the semantic integrity of the voice command to be executed, and hg is the abnormal state analysis result of the semantic compliance of the voice command to be executed.
[0024] Exemplarily, this embodiment uses a logistic function as a carrier for calculating the abnormal state of content analysis, which can compress the analysis results of any input abnormal state into a value range of 0 to 1, and can adapt to the quantitative requirements of this embodiment for the abnormal state of content analysis. Specifically, this embodiment linearly accumulates the abnormal states of semantic integrity and compliance, which can reflect the changes in the abnormal state of content analysis under the synergistic influence of the two types of abnormal states; and through an exponential function, the linear accumulation results are converted into nonlinear growth, which can reflect that the higher the degree of the two types of abnormalities, the more destructive the content analysis results of the voice instructions to be executed will be. At the same time, this embodiment uses , and further constrained the calculation results of the abnormal state of content analysis to a reasonable value range of 0 to 1, so that the calculation results can intuitively reflect the gradual change process of the abnormal state from nothing to something, from low to high, and realized the scientific quantification of the abnormal state of voice command content analysis. It can effectively reflect the comprehensive state of abnormal risks faced by content analysis of voice commands to be executed under the dual constraints of semantic integrity and compliance, and provide a reasonable and scientific quantitative basis for subsequent risk assessment of command operation intentions and command conversion.
[0025] In this embodiment, the process of constructing the instruction content parsing abnormal state analysis model in step S32 includes the following specific steps: S321. Analyze the semantic integrity abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain an analysis result of the semantic integrity abnormality state of the voice command to be executed; The calculation formula for the semantic integrity abnormal state is: ; Wherein, wz is the abnormal state of the semantic integrity of the voice command to be executed, Vs is the semantic feature vector of the voice command to be executed in the voice command data to be executed, Vst is the semantic feature vector of the semantically complete voice command of the same type as the voice command to be executed in the historical accident data, fmi is the i-th type of semantic feature vector missing in the semantically incomplete voice command of the same type as the voice command to be executed in the historical accident data, k is the number of missing semantic feature vectors in the semantically incomplete voice command of the same type as the voice command to be executed in the historical accident data, E(Vs, Vst) is the edit distance between Vs and Vst, and Len(Vst) is the vector length of Vst; I(Vs, fmi) is an indicator function, wherein when the i-th type of semantic feature vector missing in the semantically incomplete voice command of the same type as the voice command to be executed is missing in Vs, I(Vs, fmi)=1, otherwise, I(Vs, fmi)=0; wherein, the edit distance calculation formula is the existing technology and is therefore not repeated here; Exemplarily, this embodiment analyzes the abnormal state of semantic integrity of the voice instruction to be executed through semantic differences and historical missing features. Specifically, this embodiment uses semantic feature vectors as the calculation basis, measures the overall semantic difference between the voice instruction to be executed and the semantically complete voice instruction of the same type as the voice instruction to be executed by the edit distance between their semantic feature vectors, and normalizes the overall semantic difference between the two by the length of the semantic feature vector of the semantically complete voice instruction of the same type as the voice instruction to be executed; quantifies the impact of semantic missing in the semantic feature vector of the voice instruction to be executed on the execution of the instruction by using the historical missing semantic feature vector and the indicator function, and processes the operation of the historical missing semantic feature vector by using the exponential function. The semantic omissions in the instructions are transformed by an exponential function, the impact of semantic omissions on the semantic integrity of the instructions is amplified, and the discrete judgment of semantic omissions is converted into a continuous numerical impact, so that the semantic integrity abnormal state can comprehensively quantify the nonlinear impact of the overall semantic difference and semantic omissions on the semantic integrity abnormal state, and more accurately reflect the semantic integrity of the instructions; further, this embodiment also quantifies the impact of the overall semantic difference and semantic omissions on the semantic integrity abnormal state through multiplication operation; wherein the semantic feature vector of the voice instruction to be executed is constructed by semantically encoding the voice instruction to be executed and extracting key semantic elements, such as operation verbs, objects, necessary parameters, etc.; the semantic feature vector of the semantically complete voice instruction of the same type as the voice instruction to be executed is obtained by screening the semantically complete similar instructions from historical accident data and performing the same encoding; the missing semantic feature vector is formed by statistically encoding the high-frequency missing semantic elements from the historical semantically incomplete similar instructions that caused accidents; wherein the classification of similar voice instructions is to classify the operation content and field involved in the semantics of the voice instruction to be executed, and screen historical instructions with similar functions and contexts into voice instructions of the same type as the voice instruction to be executed.
[0026] S322: Analyze the semantic compliance abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain a semantic compliance abnormality state analysis result of the voice command to be executed; The calculation formula for semantic compliance exception status is: ; Where hg is the semantic compliance anomaly of the voice command to be executed, nt is the total number of semantic units in the voice command data to be executed, n is the number of non-compliant semantic units in the voice command data to be executed, Tn is the average execution time of semantically compliant voice commands of the same type as the voice command to be executed in historical accident data, and Tv is the average accident handling time of accidents caused by semantically non-compliant voice commands of the same type as the voice command to be executed in historical accident data. For example, this embodiment amplifies the impact of semantically non-compliant voice commands on the analysis of semantic compliance anomalies by performing a logarithmic operation on the ratio of the average handling time of non-compliant incidents to the average execution time of compliant commands. This converts the potentially wide-ranging duration ratio into a relatively smoothly varying value, thus avoiding imbalances in the calculation results caused by excessively large duration ratios. Furthermore, this embodiment integrates the impact of semantic compliance with the impact of incidents to quantify the actual harm caused by semantic violations on instruction conversion and execution. Semantic units are obtained by segmenting and parsing voice commands. The number of non-compliant semantic units is obtained by comparing the semantic units of semantically non-compliant voice commands in historical incident data that are similar to the voice command to be executed with the semantic unit database of historical compliant voice commands extracted from similar compliant commands, and counting the number of non-compliant semantic units. The average execution time is obtained by filtering similar semantically compliant voice commands from historical incident data and calculating the average execution time. The average incident handling time is obtained by counting similar semantically non-compliant voice commands that caused incidents and calculating the average time spent from incident to recovery during incident handling.
[0027] In this embodiment, the operation intention risk assessment model is constructed in step S4, including the following specific steps: S41, extracting and analyzing the threat level of the voice command to be executed to the currently running operation, and the abnormal state analysis result of the content parsing of the voice command to be executed; S42. Assess the risk of the operation intention of the voice command to be executed based on the analysis result of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis result of the abnormal state of the content analysis; The evaluation formula for operation intention risk is: ; Where ZY is the operation intention risk of the voice command to be executed, Cz is the analysis result of the operation threat degree caused by the voice command to be executed to the current operation, and Nr is the analysis result of the abnormal state of the content analysis of the voice command to be executed. They are the operation threat impact weight and content analysis anomaly impact weight.
[0028] For example, in the risk assessment of voice command operation intentions in this embodiment, the degree of operation threat can reflect the potential interference of command execution on the current running operation, and the abnormal state of content parsing can reflect the reliability of the semantic understanding of the command itself. The two jointly act on the operation intention risk from the two dimensions of command execution impact and command quality. Through linear weighting, the contribution of different dimensions to the operation intention risk can be clearly separated and quantified. By assigning different weights, it is possible to flexibly adapt to the difference in attention paid to operation threats and content anomalies during voice command conversion. When the voice command conversion process focuses more on the impact of command execution on existing operations, the weight of the operation threat impact can be appropriately increased; if the voice command conversion process focuses more on the quality of the command parsing itself, the weight of the content parsing anomaly impact can be appropriately increased.
[0029] In this embodiment, step S5 converts the voice command to be executed according to the operation intention risk assessment result, including the following specific steps: S51. Obtaining a risk assessment result of the operation intention of the voice command to be executed; S52: Preset an operation intention risk threshold. When the operation intention risk assessment result of the voice instruction to be executed is greater than or equal to the operation intention risk threshold, the voice instruction to be executed is rejected for conversion. When the operation intention risk assessment result of the voice instruction to be executed is less than the operation intention risk threshold, the voice instruction to be executed is converted. The setting parameters (e.g., weights and thresholds) in this embodiment are obtained by experiments conducted by those skilled in the art. The specific experimental method is as follows: obtaining historical data of multiple voice instructions to be executed, historical accident data, and corresponding current operation data; substituting the voice instruction data to be executed, historical accident data, and corresponding current operation data into each step of this embodiment to obtain the operation intention risk assessment results of the multiple voice instructions to be executed in the past; obtaining the judgment results of whether the conversion of the multiple voice instructions to be executed in the past is risky, and importing the operation intention risk assessment results of the multiple voice instructions to be executed in the past obtained from each step of this embodiment and the judgment results of whether the conversion of the multiple voice instructions to be executed in the past is risky into the fitting software, and outputting the values of the setting parameters (e.g., weights and thresholds) that meet the highest accuracy rate for judging the operation intention risk.
[0030] Example 2
[0031] like Figure 4 As shown, this embodiment provides a command conversion control system based on voice recognition, including: Data acquisition module, used to obtain voice command data to be executed, historical accident data and current operation data; An operation threat level analysis module is used to import the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; The command content parsing abnormal state analysis module is used to import the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; The operation intention risk assessment module is used to build an operation intention risk assessment model. The operation threat analysis results of the voice command to be executed on the currently running operation and the abnormal state analysis results of the content analysis are imported into the operation intention risk assessment model to assess the operation intention risk of the voice command to be executed; The command conversion control module is used to convert the voice commands to be executed based on the risk assessment results of the operation intention; The control module is used to control the operation of the data acquisition module, the operation threat level analysis module, the instruction content parsing abnormal state analysis module, the operation intention risk assessment module and the instruction conversion control module.
[0032] The above-mentioned parameters and steps for each unit module to implement corresponding functions in the command conversion control system based on voice recognition of the present invention can refer to the parameters and steps in the embodiment of the command conversion control method based on voice recognition above, and will not be repeated here.
[0033] The various embodiments of the present invention are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the IoT device and medium embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0034] The system and medium provided in the embodiments of the present invention correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.
[0035] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0036] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0037] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0038] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0039] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0040] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0041] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0042] The above are merely embodiments of the present invention and are not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A command conversion control method based on speech recognition, characterized in that: The steps include: S1. Obtaining voice command data to be executed, historical accident data, and current operation data; S2. Importing the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; S3. Importing the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; S4. Build an operation intention risk assessment model, import the analysis results of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis results of the abnormal state of the content analysis into the operation intention risk assessment model, and evaluate the operation intention risk of the voice command to be executed; S5. Convert the voice command to be executed based on the risk assessment result of the operation intention.
2. The command conversion control method based on voice recognition according to claim 1, characterized in that: The step S2 analyzes the operational threat level caused by the voice command to be executed to the currently running operation, including the following specific steps: S21, extracting the voice command data to be executed and the current operation data; S22. Based on the pending voice command data and the current running operation data, construct an operation threat level analysis model to analyze the operational threat level posed by the pending voice command to the currently running operation, and obtain an analysis result of the operational threat level posed by the pending voice command to the currently running operation; The calculation formula for the operation threat level is: ; Where Cz is the operational threat level caused by the pending voice command to the currently running operation, ct is the analysis result of the compatibility conflict threat level caused by the pending voice command to the currently running operation, and yc is the analysis result of the abnormal urgency threat level caused by the pending voice command to the currently running operation.
3. The command conversion control method based on voice recognition according to claim 2, characterized in that: The process of constructing the operational threat level analysis model in step S22 includes the following specific steps: S221. Analyzing the compatibility conflict threat level caused by the voice command to be executed on the currently running operation based on the voice command data to be executed and the currently running operation data, thereby obtaining a compatibility conflict threat level analysis result caused by the voice command to be executed on the currently running operation; S222. Based on the voice instruction data to be executed and the current operation data, analyze the abnormal urgency threat degree caused by the voice instruction to be executed to the currently running operation, and obtain an analysis result of the abnormal urgency threat degree caused by the voice instruction to be executed to the currently running operation.
4. The command conversion control method based on voice recognition according to claim 3, characterized in that: The step S3 analyzes the abnormal state of the content analysis of the voice command to be executed, including the following specific steps: S31, extracting voice command data to be executed and historical accident data; S32. Based on the voice command data to be executed and the historical accident data, a command content parsing abnormal state analysis model is constructed, and the abnormal state of the content parsing of the voice command to be executed is analyzed to obtain an analysis result of the abnormal state of the content parsing of the voice command to be executed; The calculation formula for content parsing abnormal status is: ; Where Nr is the abnormal state of content analysis of the voice command to be executed, wz is the abnormal state analysis result of the semantic integrity of the voice command to be executed, and hg is the abnormal state analysis result of the semantic compliance of the voice command to be executed.
5. The command conversion control method based on voice recognition according to claim 4, characterized in that: The construction process of the instruction content parsing abnormal state analysis model in step S32 includes the following specific steps: S321. Analyze the semantic integrity abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain an analysis result of the semantic integrity abnormality state of the voice command to be executed; S322: Analyze the semantic compliance abnormality state of the voice command to be executed based on the voice command data to be executed and the historical accident data to obtain an analysis result of the semantic compliance abnormality state of the voice command to be executed.
6. The command conversion control method based on speech recognition according to claim 5, characterized in that: The step S4 constructs an operation intention risk assessment model, including the following specific steps: S41, extracting and analyzing the threat level of the voice command to be executed to the currently running operation, and the abnormal state analysis result of the content parsing of the voice command to be executed; S42. Assess the risk of the operation intention of the voice command to be executed based on the analysis result of the operational threat level caused by the voice command to be executed to the currently running operation and the analysis result of the abnormal state of the content analysis; The evaluation formula for operation intention risk is: ; Where ZY is the operation intention risk of the voice command to be executed, Cz is the analysis result of the operation threat degree caused by the voice command to be executed to the current operation, and Nr is the analysis result of the abnormal state of the content analysis of the voice command to be executed. They are the operation threat impact weight and content analysis anomaly impact weight.
7. The command conversion control method based on voice recognition according to claim 6, characterized in that: In step S5, the voice command to be executed is converted according to the operation intention risk assessment result, including the following specific steps: S51. Obtaining a risk assessment result of the operation intention of the voice command to be executed; S52. Preset an operation intention risk threshold. When the operation intention risk assessment result of the voice instruction to be executed is greater than or equal to the operation intention risk threshold, refuse to convert the voice instruction to be executed; when the operation intention risk assessment result of the voice instruction to be executed is less than the operation intention risk threshold, convert the voice instruction to be executed.
8. A command conversion control system based on speech recognition, which is implemented based on the command conversion control method based on speech recognition according to any one of claims 1 to 7, characterized in that: The system comprises: Data acquisition module, used to obtain voice command data to be executed, historical accident data and current operation data; An operation threat level analysis module is used to import the voice command data to be executed and the current operation data into the operation threat level analysis model to analyze the operational threat level caused by the voice command to be executed to the current operation; The command content parsing abnormal state analysis module is used to import the voice command data to be executed and the historical accident data into the command content parsing abnormal state analysis model to analyze the abnormal state of the content parsing of the voice command to be executed; The operation intention risk assessment module is used to build an operation intention risk assessment model. The operation threat analysis results of the voice command to be executed on the currently running operation and the abnormal state analysis results of the content analysis are imported into the operation intention risk assessment model to assess the operation intention risk of the voice command to be executed; The command conversion control module is used to convert the voice commands to be executed based on the risk assessment results of the operation intention; A control module is used to control the operation of the data acquisition module, the operation threat level analysis module, the instruction content parsing abnormal state analysis module, the operation intention risk assessment module and the instruction conversion control module.
Citation Information
Patent Citations
Instruction management method and device, equipment and computer readable storage medium
CN112667290A
Operation instruction processing method and device and electronic equipment
CN116150742A
Operation system operation and maintenance auxiliary method and system based on deep reinforcement learning
CN117492807A
Instruction monitoring method and device, electronic equipment and medium
CN117787707A
Intelligent equipment anti-error identification and automatic auditing system
CN119477204A
Cited By
Voice control method and system for wireless handheld ultrasonic equipment
CN121506133A