Voice control method and system of virtual platform, and storage medium

By adopting a voice control method based on the dual cognitive mechanism of ‘conditioned reflex + deep thinking’ on the virtual platform, the problem of response speed and complex command processing problems in the prior art is solved, efficient voice control is achieved, and the error operation rate is reduced.

CN120183402AActive Publication Date: 2025-06-20CHENGDU TME SOFTWARE
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202510654223.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The prior art methods used for voice control in virtual platforms cannot take into account the difficulties of response speed and complex command processing, resulting in low efficiency, especially in multi-language environments with high error operation rate.

Method used

The voice control method based on the dual cognitive mechanism of ‘conditioned reflection + deep thinking’ is adopted. By denoising the initial voice command and identifying the start and end points, it is mapped to a predefined atomic operation instruction set, and atomic operation instructions or instruction sequences are generated according to the complexity of the instruction, and the verification is carried out in combination with the current content and natural laws of the virtual platform, and the verification passed instructions are finally executed.

Benefits of technology

It improves the response speed and accuracy of voice control, reduces the error operation rate, and enhances the overall efficiency of voice control on the virtual platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183402A_ABST
    Figure CN120183402A_ABST
Patent Text Reader

Abstract

The invention provides a voice control method and system for a virtual platform and a storage medium, and the method comprises the steps: carrying out the noise reduction and starting and ending point recognition of an obtained initial voice instruction, and obtaining a processed voice instruction; mapping a text converted by the processed voice instruction to a predefined atomic operation instruction set to obtain an initialization instruction; generating an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction; the atomic operation instruction and the atomic operation instruction sequence are instructions capable of being directly executed; verifying the rationality of the atomic operation instruction or the atomic operation instruction sequence in combination with the currently virtualized content of the virtual platform and the natural law; and executing the checked atomic operation instruction or atomic operation instruction sequence. According to the method, the problem that the response speed and complex instruction processing cannot be both considered in the current common voice control method can be solved, so that the efficiency of performing voice control on the virtual platform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of virtual platform control, and more particularly, to a voice control method, system, and storage medium for a virtual platform. Background Art

[0002] With the development of the times, technical means such as WebGL 3D rendering and physical engine simulation have realized the digital reconstruction of experimental processes in disciplines such as chemistry, biology, and physics. In the evolution of virtual experiment interaction methods, voice control technology has gradually become a key breakthrough for improving operation efficiency.

[0003] However, the commonly used voice control methods currently expose serious adaptability problems in virtual platforms. Taking a typical chemical titration experiment as an example, users need to perform multiple click operations, with a high accidental touch rate and a relatively long average operation time. At the same time, in a multilingual environment, the misoperation rate caused by semantic differences in operations is increasing steadily.

[0004] That is to say, the current method for voice control of virtual platforms is not efficient enough. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a voice control method, system, and storage medium for a virtual platform. Based on the dual cognitive mechanism of "conditioned reflex + deep thinking", this method can solve the problem that the commonly used voice control methods currently cannot balance response speed and complex instruction processing during the control of virtual platforms, thereby improving the efficiency of voice control of virtual platforms.

[0006] In a first aspect, the embodiments of this application provide a voice control method for a virtual platform, including: performing noise reduction and start / end point recognition on the obtained initial voice command to obtain a processed voice command; mapping the text converted from the processed voice command to a predefined atomic operation instruction set to obtain an initialization instruction; generating an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction; where the atomic operation instruction and the atomic operation instruction sequence are respectively instructions that can be directly executed; combining the content currently virtualized by the virtual platform and the natural laws to verify the rationality of the atomic operation instruction or the atomic operation instruction sequence; and executing the verified atomic operation instruction or atomic operation instruction sequence.

[0007] The voice control method of the above virtual platform performs noise reduction and start / end point recognition on the obtained initial voice command, maps the converted text to a predefined atomic operation instruction set to obtain an initialized voice command, and performs corresponding processing based on the complexity of the initialized voice command to generate an atomic operation instruction or an atomic operation instruction sequence that can be directly executed. Before executing the generated atomic operation instruction or atomic operation instruction sequence, its rationality is verified, which improves the processing speed and accuracy of the voice, thereby improving the efficiency of voice control of the virtual platform.

[0008] In combination with the first aspect, optionally, the noise reduction and start / end point recognition of the obtained initial voice command to obtain a processed voice command includes: representing the initial voice command as a time-domain signal; and using spectral subtraction to remove the noise component in the initial voice command; wherein, the implementation formula of the spectral subtraction is: ; In the formula, f represents the frequency, X clean (f) represents the pure voice signal after removing the noise component from the initial voice command, X(f) represents the frequency-domain signal after Fourier transform of the time-domain signal, N(f) represents a preset room environment noise template; α is a subtraction factor, β is a spectral lower limit protection coefficient, represents the phase information of the initial voice command.

[0009] The voice control method of the above virtual platform performs noise reduction processing on the initial voice command by using spectral subtraction, removes the noise component from the noisy voice spectrum, ensures the pure voice characteristics, thereby obtaining a clearer and more easily recognizable voice signal, and also facilitates the subsequent accurate recognition of the start / end point. Finally, it further improves the processing speed and accuracy of the voice, and correspondingly further improves the efficiency of voice control of the virtual platform.

[0010] In combination with the first aspect, optionally, the noise reduction and start / end point recognition of the obtained initial voice command to obtain a processed voice command further includes: converting the spectrum after inverse Fourier transform of the frequency-domain signal into a time-domain signal; judging whether the number of consecutive frames satisfying and in the time-domain signal exceeds a first frame number threshold; if it is determined to be satisfied, taking the first frame in the consecutive frames as the start point of the voice command; judging whether the number of consecutive frames satisfying and in the time-domain signal exceeds a second frame number threshold; if it is determined to be satisfied, taking the last frame in the consecutive frames as the end point of the voice command; wherein, E n and Z nCalculated by the following formulas respectively: ; ; In the formula, E n represents the short-time energy of the nth frame in the time-domain signal, L represents the frame length, x(k) represents the amplitude of the kth sampling point, Z n represents the zero-crossing rate of the nth frame in the time-domain signal, I represents a judgment function for judging whether it holds, μE and μZ respectively represent the mean of the short-time energy and the mean of the zero-crossing rate of the noise component; σE and σZ represent the standard deviation of the short-time energy and the standard deviation of the zero-crossing rate of the noise component.

[0011] The voice control method of the above virtual platform calculates the short-time energy and zero-crossing rate of the initial voice command converted into a time-domain signal, and determines the start and end points by judging whether the short-time energy and zero-crossing rate simultaneously satisfy the relationships corresponding to the mean of the short-time energy and zero-crossing rate of the noise component, the standard deviation of the short-time energy and zero-crossing rate of the noise component. It more effectively distinguishes the voice signal from the background noise, reduces false triggers, improves the accuracy of voice recognition, and ultimately further improves the processing speed and accuracy of the voice, and correspondingly further improves the efficiency of voice control of the virtual platform.

[0012] Combined with the first aspect, optionally, mapping the text converted from the processed voice command to a predefined atomic operation instruction set to obtain an initialization instruction includes: performing frame segmentation on the processed voice command and inputting it into a pre-constructed semantic feature extraction module to generate a phoneme sequence; inputting the phoneme sequence into a pre-constructed semantic feature alignment module to generate a semantic vector; and performing a normalization mapping on the semantic vector to obtain the initialization instruction; where the formula for the normalization mapping is: ; In the formula, z 标 represents the initialization instruction, h CLS is the semantic vector, z is the predefined atomic operation instruction set, z i is an element in the atomic operation instruction set, and E(z i ) represents the pre-trained instruction embedding vector.

[0013] The voice control method of the above virtual platform realizes the conversion of the processed voice command into a semantic vector by performing frame splitting and semantic feature alignment on the processed voice command. Then, the semantic vector is subjected to a normalization mapping to generate a structured and initialized command that can be directly used for subsequent processing, ensuring the accuracy and executability of the command, thereby further improving the processing speed and accuracy of the voice, and correspondingly further improving the efficiency of voice control of the virtual platform.

[0014] In combination with the first aspect, optionally, generating an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction includes: calculating the instruction complexity of the initialization instruction by a reflexive neural network; wherein, the reflexive neural network includes a fast response channel and a complex instruction processing channel; the calculation formula of the instruction complexity is as follows: ; In the formula, lexical_complexity represents the lexical complexity of the initialization instruction, syntactic_depth represents the syntactic depth of the initialization instruction, and Score represents the instruction complexity; determining whether the instruction complexity exceeds a complexity threshold; if it is determined to exceed, the complex instruction processing channel receives the initialization instruction and the cached historical operation information, and performs an embedding layer transformation on the initialization instruction to output an embedding vector; calculating the relative position encoding of the initialization instruction by using a position encoding formula; wherein, the position encoding formula is: ; In the formula, k represents a dimension index, d represents an embedding dimension, and i and j respectively represent the sequence positions of two elements in the initialization instruction; fusing the embedding vector with the relative position encoding to obtain a fused embedding vector; processing the fused embedding vector by using a multi-head attention mechanism to obtain a context-aware vector; and combining the context-aware vector with the historical operation information to generate the atomic operation instruction or the atomic operation instruction sequence.

[0015] The voice control method of the above virtual platform starts from the input semantic vector sequence and historical operation information, and through the processing of relative position encoding and multi-head attention mechanism, enables the complex instruction processing channel to effectively capture the position information in the sequence and integrate it into the understanding and processing of complex instructions. This helps to more accurately parse instructions with complex syntactic structures and long-distance dependencies, and generate accurate operation instructions in combination with historical context. That is, it improves the processing speed and accuracy of the voice, thereby improving the efficiency of voice control of the virtual platform.

[0016] In combination with the first aspect, optionally, generating atomic operation instructions or an atomic operation instruction sequence according to the complexity of the initialization instruction further includes: if it is determined that the limit is not exceeded, performing a convolution operation on the Mel-frequency cepstral coefficients of the initialization instruction by the fast response channel to capture key acoustic patterns; performing global pooling dimensionality reduction on the initialization instruction after the convolution operation, and outputting the probability distribution of the elements in the atomic operation instruction set by using an activation function; and generating the atomic operation instruction or the atomic operation instruction sequence according to the probability distribution.

[0017] The above voice control method for the virtual platform realizes the fast processing of simple instructions, the efficient recognition and fast response of simple instructions by performing convolution operation on the Mel-frequency cepstral coefficients of the initialization instruction, global pooling dimensionality reduction and output of the probability distribution by the activation function, and combining the probability distribution to output the atomic operation instruction or the atomic operation instruction sequence. Finally, it also improves the processing speed and accuracy of the voice, thereby improving the efficiency of voice control of the virtual platform.

[0018] In combination with the first aspect, optionally, verifying the rationality of the atomic operation instruction or the atomic operation instruction sequence by combining the content currently virtualized by the virtual platform and the natural laws includes: constructing a four-dimensional state tensor for describing the virtual platform screen; where the tensor formula is: ; In the formula, x, y, z represent the scene coordinate system of the virtual platform screen, t represents the time stamp, v represents the state of the operation object corresponding to the atomic operation instruction or the atomic operation instruction sequence, and c p represents the operation stage divided for the operation process; r represents the constraint rule; Extracting the image features of the virtual platform screen to obtain an image feature vector; performing linear projection on the atomic operation instruction or the atomic operation instruction sequence to obtain a projected voice feature vector; combining the image feature vector and the projected voice feature vector to perform cross-modal attention calculation to obtain an attention weight distribution; where the formula for cross-modal attention calculation is: ; In the formula, Q i is the linear projection of the voice feature vector, K j is the linear projection of the image feature vector, and d is the dimension of the voice feature vector and the image feature vector; verifying the rationality of the atomic operation instruction or the atomic operation instruction sequence by combining the attention weight distribution and the four-dimensional state tensor.

[0019] The above-mentioned voice control method for the virtual platform can determine the specific position or object in the virtual experimental platform referred to by the current voice command (atomic operation command or atomic operation command sequence) through the attention weight distribution, so as to verify the rationality of the atomic operation command or atomic operation command sequence, ensuring the rationality, feasibility, and security of the command in the experimental environment, etc. This further improves the processing speed and accuracy of the voice, and thus improves the efficiency of voice control for the virtual platform.

[0020] In combination with the first aspect, optionally, verifying the rationality of the atomic operation command or atomic operation command sequence by combining the attention weight distribution and the four-dimensional state tensor includes: verifying the atomic operation command or atomic operation command sequence according to the natural laws by combining the attention weight distribution and the four-dimensional state tensor; and verifying whether the operation object corresponding to the atomic operation command or atomic operation command sequence is consistent with the actual object in the virtual platform screen by comparing the object positioning heat map corresponding to the attention weight distribution with the four-dimensional state tensor.

[0021] The above-mentioned voice control method for the virtual platform verifies the command by comparing the relevant area in the object positioning heat map with the coordinate system in the four-dimensional state tensor to confirm whether the operation object in the command is consistent with the actual object in the virtual platform, so as to verify the command. This further ensures the rationality, feasibility, and security of the command in the experimental environment, etc. This further improves the processing speed and accuracy of the voice, and thus improves the efficiency of voice control for the virtual platform.

[0022] In a second aspect, an embodiment of the present application further provides a voice control system for a virtual platform, including: a voice input layer for denoising and start / stop point recognition of the acquired initial voice command to obtain a processed voice command; a multi-language processing layer for mapping the text converted from the processed voice command to a predefined atomic operation command set to obtain an initialization command; a reflexive neural network for generating an atomic operation command or atomic operation command sequence according to the complexity of the initialization command, where the atomic operation command and the atomic operation command sequence are respectively commands that can be directly executed; a spatio-temporal perception matrix for verifying the rationality of the atomic operation command or atomic operation command sequence by combining the content currently virtualized by the virtual platform and the natural laws; and an execution layer for executing the verified atomic operation command or atomic operation command sequence.

[0023] The above-mentioned voice control system device for the virtual platform has the same beneficial effects as the voice control method for the virtual platform provided in the first aspect or any optional implementation manner of the first aspect, and will not be elaborated here.

[0024] In a third aspect, an embodiment of the present application further provides a storage medium, which includes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the method described above.

[0025] The above storage medium has the same beneficial effects as the voice control method of the virtual platform provided in the first aspect or any optional implementation manner of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is a flowchart of the voice control method of the virtual platform provided by the embodiment of the present application; Figure 2 is a specific flowchart of step S110 in the voice control method of the virtual platform provided by the embodiment of the present application; Figure 3 is a specific flowchart of step S120 in the voice control method of the virtual platform provided by the embodiment of the present application; Figure 4 is a specific flowchart of step S130 in the voice control method of the virtual platform provided by the embodiment of the present application; Figure 5 is a specific flowchart of step S140 in the voice control method of the virtual platform provided by the embodiment of the present application; Figure 6 is a specific flowchart of step S145 in the voice control method of the virtual platform provided by the embodiment of the present application; Figure 7 is a schematic diagram of the voice control system of the virtual platform provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The embodiments of the technical solutions of the present application will be described in detail below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and therefore are only examples and cannot be used to limit the protection scope of the present application.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0030] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality" means more than two unless otherwise specifically defined.

[0031] Please refer to Figure 1 , Figure 1 is a flowchart of the voice control method for the virtual platform provided by the embodiments of this application. The voice control method for the virtual platform provided by the embodiments of this application may include: Step S110: Perform noise reduction and start / end point recognition on the obtained initial voice command to obtain a processed voice command.

[0032] In the above step S110, voice data can be obtained first through an audio input device such as a microphone. Specifically, a multi-microphone array (4 channels, 120° pickup angle), a sampling rate of 48 kHz, and a quantization accuracy of 24 bits can be used, and a high-frequency noise can be eliminated through an 8th-order Butterworth low-pass filter (cutoff frequency 20 kHz). After performing noise reduction and other preliminary processing on the initial voice command, since during the process of voice control, the time period during which audio can be collected is usually not limited to the time period when the user issues the control voice, but in addition to this time period, there is also a lot of complex background noise. Therefore, the start point and end point of the initial voice command can be recognized to exclude the interference brought by the audio outside the time period when the user issues the control voice.

[0033] Regarding the recognition of the start / end point, specifically, the double-threshold detection algorithm can be used to calculate the short-time energy and the zero-crossing rate to accurately judge the start and end points of the voice. Even in a complex noise environment such as the sound of an air conditioner or the sound of keyboard typing, the user's voice command can be accurately recognized.

[0034] Step S120: Map the text converted from the processed voice command to a predefined atomic operation instruction set to obtain an initialization instruction.

[0035] In the above step S120, specifically, cross - language semantic alignment technology can be used to generate semantic vectors and map the text to a predefined atomic operation instruction set to obtain initialization instructions. Among them, atomic operation instructions and atomic operation instruction sequences are instructions that can be directly executed. However, the initialization instructions obtained at this stage cannot be directly executed. Taking the evolution of virtual experiments as an example, the instructions obtained at this stage cannot be directly executed because they have not been combined with the experimental state, the experimental resource state has not been considered, the parameter integrity is insufficient, the parameter rationality is in doubt, and the timing relationship is not clear, etc.

[0036] Step S130: Generate atomic operation instructions or atomic operation instruction sequences according to the complexity of the initialization instructions.

[0037] In the above step S130, atomic operation instructions and atomic operation instruction sequences are instructions that can be directly executed. According to the complexity of the initialization instructions, different processing channels can be set. For example, if the complexity of the initialization instructions is divided into three levels, then three channels can be correspondingly set to process initialization instructions of different levels of complexity respectively. Another example: if the complexity of the initialization instructions is divided into two levels, then two channels can be correspondingly set to process initialization instructions of different levels of complexity respectively. Of course, those skilled in the art can divide the complexity of the initialization instructions into corresponding numbers of levels according to actual application requirements and set the corresponding numbers of channels. Different algorithms can be used in these channels to handle initialization instructions of different complexities. The embodiments of the present application do not make specific limitations on this.

[0038] Exemplarily, the initialization instructions are divided into two levels: simple and complex. For simple and high - frequency instructions, such as "start titration" or "stop", they can be quickly processed through the corresponding channel. The network structure of this channel includes an input layer, a convolutional layer, a global pooling layer, and an output layer, and can quickly generate the probability of atomic operation instructions. It can directly trigger the predefined "start titration" atomic operation instruction. This operation instruction corresponds to a series of simple and fixed system actions, such as initializing the burette position, setting the initial volume, starting the titration timer, etc. These instructions are relatively simple and concise, that is, atomic operation instructions.

[0039] For complex instructions such as "After titrating to a pH value of 7, record the current volume and calculate the concentration", another channel can be correspondingly enabled. Combining previous experimental operations (for example, steps such as the initial preparation of acid-base solutions and the start of titration may have been carried out before) and the current experimental state (such as the pH value of the current solution may be higher or lower than 7, information such as the volume of the titrated solution, etc.), understand the complete logic of this complex instruction. Decompose the complex instruction into a series of specific operation steps. First, continuously monitor the change in the pH value of the solution; secondly, pause the titration when the pH value reaches 7 and record the volume of the solution in the burette at this time; finally, calculate the concentration of the sodium hydroxide solution using chemical formulas based on the initial concentration and titration volume of the known acid-base solution. Generate a sequence of multiple atomic operation instructions such as "monitor pH value", "compare pH value with target value", "pause titration", "record volume", "apply concentration calculation formula", etc., and determine the execution order of these operations. These instructions are relatively complex, that is, an atomic operation instruction sequence composed of multiple atomic operation instructions.

[0040] Step S140: Combine the content currently virtualized by the virtual platform and the laws of nature to verify the rationality of the atomic operation instruction or the atomic operation instruction sequence.

[0041] In the above step S140, the feasibility of the instruction in the current experimental environment can be verified by constructing a four-dimensional state tensor, combining the experimental scenario, time, object state, and experimental stage, and ensuring that the voice instruction is accurately matched with the objects and operations in the experimental scenario.

[0042] Continuing with the example of the evolution of a virtual experiment, specifically, first, check whether the resources involved in the experiment (such as solutions, instruments) are available, and whether the state of the experimental object meets the requirements of the instruction. For example, if the instruction is "Heat the solution to 80 °C", it will be verified whether the heating equipment on the virtual experimental bench is ready, whether the solution container is placed correctly, etc. Then, ensure that the atomic operation instruction or the atomic operation instruction sequence conforms to the experimental laws of physics, chemistry, etc. For example, in a chemical experiment, if the atomic operation instruction or the atomic operation instruction sequence is "Mix 100 ml of water with 500 ml of hydrochloric acid", it will be checked according to chemical common sense and experimental settings whether the volume of the mixed solution is reasonable and whether there will be a violent reaction, etc. Verify whether the instruction conforms to the current stage of the experiment. For example, in a chemical titration experiment, if the current stage is the "preparation stage" and the instruction is "Record the data at the titration end point", it will be determined that this instruction does not match the current stage and needs to be further confirmed or corrected.

[0043] Step S150: Execute the verified atomic operation instruction or atomic operation instruction sequence.

[0044] In the above step S150, the virtual platform can combine 3D rendering technology (such as WebGL), a physics engine, and interaction design to simulate the visual and operational experience of a real experiment. By interacting with the virtual platform, the user can calculate and update the experimental results in real time according to the user's operations, and provide visual feedback.

[0045] Exemplarily, in a virtual experiment of chemical titration, the user can drag a virtual dropper to drip an acid or base solution into a virtual conical flask. The system will calculate the change in the pH value of the solution in real time according to the dosage and speed of the user's operation, and display the color change or pH curve on the screen.

[0046] In the above implementation process, by performing noise reduction and start / end point recognition on the obtained initial voice command, and mapping the text converted therefrom to a predefined atomic operation instruction set, an initialized voice command is obtained, and corresponding processing is performed based on the complexity of the initialized voice command to generate an atomic operation instruction or an atomic operation instruction sequence that can be directly executed. Before executing the generated atomic operation instruction or atomic operation instruction sequence, its rationality is verified, which improves the processing speed and accuracy of the voice, thereby improving the efficiency of voice control of the virtual platform.

[0047] Please refer to Figure 2 , Figure 2 which is the specific flowchart of step S110 in the voice control method of the virtual platform provided by the embodiments of the present application. In some alternative embodiments, step S110 may include: Step S111: Represent the initial voice command as a time-domain signal.

[0048] In the above step S111, the initial voice command can be represented as a time-domain waveform signal x(t), where t is a time variable, and the unit can be milliseconds.

[0049] Step S112: Use spectral subtraction to remove the noise component in the initial voice command.

[0050] In the above step S112, the implementation formula of spectral subtraction is: ; where f represents frequency, X clean (f) represents the clean voice signal after removing the noise component from the initial voice command, X(f) represents the frequency-domain signal obtained by Fourier-transforming the time-domain signal, N(f) represents a preset room environment noise template. α is a subtraction factor, β is a spectral lower limit protection coefficient, represents the phase information of the initial voice command. Spectral subtraction can ensure the accuracy of subsequent recognition processing of the initial voice command.

[0051] In the above implementation process, the initial voice command is denoised by using spectral subtraction, so that the noise component is removed from the noisy speech spectrum, ensuring pure speech features, thereby obtaining a clearer and more easily recognizable speech signal, which also facilitates the subsequent accurate recognition of the start and end points. Finally, the processing speed and accuracy of speech are further improved, and correspondingly, the efficiency of voice control of the virtual platform is further improved.

[0052] Please continue to refer to Figure 2 , in some alternative embodiments, step S110 may further include: Step S113: Convert the spectrum obtained by inverse Fourier transform of the frequency-domain signal into a time-domain signal.

[0053] Step S114: Determine whether in the time-domain signal, the number of consecutive frames satisfying and exceeds the first frame number threshold.

[0054] In the above step S114, E n and Z n are calculated by the following formulas respectively: ; ; In the formula, E n represents the short-time energy of the nth frame in the time-domain signal, L represents the frame length, x(k) represents the amplitude of the kth sampling point, Z n represents the zero-crossing rate of the nth frame in the time-domain signal, I represents a judgment function for judging whether holds, μE and μZ respectively represent the mean of the short-time energy and the mean of the zero-crossing rate of the noise component; represents the standard deviation of the short-time energy and the standard deviation of the zero-crossing rate of the noise component.

[0055] If it is determined to be satisfied, then execute step S115: Take the first frame in the consecutive frames as the starting point of the voice command.

[0056] In the above step S115, in order to avoid misjudgment caused by short-term fluctuations, it can be detected whether multiple consecutive frames all meet the above conditions. Exemplarily, if 4 consecutive frames (exceeding the first frame number threshold "3") all meet and , it is considered that the speech signal starts, that is, the starting point. Of course, those skilled in the art can also determine the first frame number threshold according to actual needs.

[0057] Step S116: Determine whether in the time-domain signal, the number of consecutive frames satisfying and exceeds the second frame number threshold.

[0058] In the above step S116, similarly, to avoid misjudgment caused by short-term fluctuations, it is possible to detect whether the above conditions are satisfied for multiple consecutive frames. Exemplarily, if four consecutive frames (exceeding the first frame number threshold of "3") all satisfy and , it is considered that the voice signal ends, that is, the termination point. Of course, those skilled in the art can also determine the second frame number threshold according to actual needs.

[0059] If it is determined to be satisfied, step S117 is executed: using the last frame in the consecutive frame numbers as the termination point of the voice command.

[0060] In the above step S117, E n and Z n are respectively calculated by the following formulas: ; ; In the formula, E n represents the short-time energy of the nth frame in the time-domain signal, L represents the frame length, x(k) represents the amplitude of the kth sampling point, and Z n represents the zero-crossing rate of the nth frame in the time-domain signal, I represents a judgment function for judging whether it holds, μE and μZ respectively represent the mean of the short-time energy of the noise component and the mean of the zero-crossing rate; σE and σZ represent the standard deviation of the short-time energy of the noise component and the standard deviation of the zero-crossing rate.

[0061] In the above implementation process, by calculating the short-time energy and zero-crossing rate of the initial voice command converted into the time-domain signal, and by judging whether the short-time energy and zero-crossing rate simultaneously satisfy the relationships corresponding to the mean of the short-time energy of the noise component and the zero-crossing rate, the standard deviation of the short-time energy of the noise component and the standard deviation of the zero-crossing rate, the start and end points are determined. It more effectively distinguishes the voice signal from the background noise, reduces mis-triggering, improves the accuracy of speech recognition, and ultimately further improves the processing speed and accuracy of speech, and correspondingly further improves the efficiency of voice control of the virtual platform.

[0062] Please refer to Figure 3 , Figure 3 which is the specific flowchart of step S120 in the voice control method of the virtual platform provided by the embodiment of the present application. In some alternative embodiments, step S120 may include: Step S121: Perform frame splitting on the processed voice command and input it into a pre-constructed semantic feature extraction module to generate a phoneme sequence.

[0063] In the above step S121, the processed speech command can be processed based on the constructed WeNet multilingual acoustic model. The network structure of this model can specifically be: Encoder: 8 layers of Conformer (Convolutional Enhanced Transformer), each layer contains 4 heads of self-attention (head dimension 256), the convolutional kernel size is 15, and the gated linear unit (GLU) is activated; Decoder: 2 layers of Transformer, trained jointly using CTC / Attention.

[0064] Exemplarily, the processed speech command is framed into 128-dimensional Mel filter bank features (frame length 25ms, frame shift 10ms). After inputting it into the above model, the model outputs a phoneme sequence , where . And it can support shared encoders for 12 languages.

[0065] Step S122: Input the phoneme sequence into the pre-constructed semantic feature alignment module to generate a semantic vector.

[0066] In the above step S122, the XLM-R model can specifically be used as the semantic feature alignment module for cross-language semantic alignment to generate a semantic vector: ; where h CLS represents the semantic vector corresponding to the position of the CLS token, XLM-R base represents the XLM-R base model, s represents the input text sequence (such as the Chinese "Centrifuge for 5 minutes"), CLS represents the token, and the vector representation of the CLS token is usually used as the semantic representation of the entire sequence.

[0067] Step S123: Perform a normalization mapping on the semantic vector to obtain an initial instruction.

[0068] In the above step S123, the formula for the normalization mapping is: ; where z 标 represents the initial instruction, h CLS is the semantic vector, z is the predefined atomic operation instruction set, z i is an element in the atomic operation instruction set, and E(z i ) represents the pre-trained instruction embedding vector.

[0069] Exemplarily: Input statement: "Centrifuge for 5 minutes"; Output {action: centrifuge, duration: 300}.

[0070] In the above implementation process, through frame processing and semantic feature alignment of the processed voice command, the conversion of the processed voice command into a semantic vector is achieved. Then, the semantic vector is subjected to a normalization mapping to generate a structured and initialized command that can be directly used for subsequent processing, ensuring the accuracy and executability of the command, thereby further improving the processing speed and accuracy of the voice, and correspondingly further improving the efficiency of voice control of the virtual platform.

[0071] Please refer to Figure 4 , Figure 4 is the specific flowchart of step S130 in the voice control method of the virtual platform provided by the embodiment of the present application. In some optional implementation manners, step S130 may include: Step S131: Calculate the command complexity of the initialized command by a recurrent neural network.

[0072] In the above step S131, the recurrent neural network may include a fast response channel and a complex command processing channel. The fast response channel may specifically be a reflection channel, and the complex command processing channel may specifically be a cognitive channel. The network structure of the reflection channel may specifically be as shown in Table 1.

[0073] Table 1 Hierarchy Operation Output Dimension Function Input Layer MFCC Feature Input 128×256 Extract Speech Time-Frequency Characteristics Convolutional Layer 1 3×3 Convolution, stride = 2, ReLU Activation 64×128×128 Capture Local Acoustic Patterns Convolutional Layer 2 5×5 Convolution, stride = 1, ReLU Activation 32×124×124 Enhance Feature Abstraction Ability Global Pooling Layer Average Pooling 32×1×1 Reduce Dimension and Retain Key Information Output Layer Sigmoid Activation 286-Dimensional Probability Distribution Generate Probability of Atomic Operation Instructions The calculation formula of the command complexity is as follows: ; In the formula, lexical_complexity represents the lexical complexity of the initialized command, syntactic_depth represents the syntactic depth of the initialized command, and Score represents the command complexity.

[0074] Step S132: Determine whether the command complexity exceeds the complexity threshold.

[0075] In the above step S132, exemplarily, the complexity threshold may be 6.5, then it can be determined whether Score > 0.65 holds. Of course, those skilled in the art can determine other values of the complexity threshold according to the actual application situation.

[0076] If it is determined that it exceeds, then execute step S133: The complex command processing channel receives the initialized command and the cached historical operation information, and performs an embedding layer transformation on the initialized command to output an embedding vector.

[0077] In the above step S133, if it is determined that it exceeds, it means that the complexity of the initialization instruction is relatively high, and it can be processed by the complex instruction processing channel. Specifically, it can be implemented as follows: first, receive the initialization instruction and the cached historical operation information. The core component of the cognitive channel can be a memory unit, which can cache the hidden states of the most recent 20 instructions. Convert each element (such as a word or a sub-word unit, etc.) in the initialization instruction into a vector representation of a fixed dimension. Exemplarily, assume that the input is a tokenized instruction sequence , each token w i is converted into a vector e i through the embedding layer, forming an embedding matrix .

[0078] Step S134: Calculate the relative position encoding of the initialization instruction using the position encoding formula.

[0079] In the above step S134, the position encoding formula is as follows: ; In the formula, k represents the dimension index, d represents the embedding dimension, and i and j respectively represent the sequence positions of two elements in the initialization instruction. The embedding dimension can specifically be 512. Specifically, i and j in the position encoding formula respectively represent the position indices of two elements in the sequence: i represents the position of the current element in the sequence; j represents the position of another element in the sequence.

[0080] The goal of the relative position encoding is to calculate the relative distance (i - j) between these two positions, and then map this relative distance to a high-dimensional space through a sine or cosine function to generate the corresponding position encoding. Through the relative position encoding, the relative position relationship between elements can be reflected.

[0081] Step S135: Fuse the embedding vector with the relative position encoding to obtain a fused embedding vector.

[0082] In the above step S135, the fusion method can specifically be addition. For example, the embedding vector is e i , the relative position encoding is R ij , and the fused vector is e i ' = e i + R ij .

[0083] Step S136: Process the fused embedding vector using the multi-head attention mechanism to obtain a context-aware vector.

[0084] In the above step S136, the multi-head attention mechanism calculates in parallel through multiple attention heads to capture the associations between different elements in the sequence. Each attention head calculates query (Query), key (Key), and value (Value) vectors, and obtains the output through weighted summation of the attention scores. The multi-head attention mechanism outputs a new set of vector representations, that is, context-aware vectors. These vectors fuse the semantic information, relative position information, and association information of the input elements. These vector representations not only contain the semantic information and relative position information of the input elements, but also capture the association information between the elements.

[0085] Among them, the formula of the multi-head attention mechanism can be: ; In the formula, each attention head is defined as: , where Q is the query vector, representing the features of the position or element that needs to be focused on currently. K is the key vector, representing the features of other positions or elements, which is used to match with the query vector. V is the value vector, representing the actual information content to be focused on. is the output weight matrix, which is used to linearly transform the concatenated vector into the final output space. Concat represents the concatenation operation, which is a function that concatenates the outputs of multiple attention heads into a vector. W i Q is the query weight matrix of the i-th attention head, which is used to linearly transform the input query vector into the i-th subspace. W i K is the key weight matrix of the i-th attention head, which is used to linearly transform the input key vector into the i-th subspace. W i V is the value weight matrix of the i-th attention head, which is used to linearly transform the input value vector into the i-th subspace. Attention is a function of the process of calculating the attention scores and weighted summing the value vectors.

[0086] Step S137: Combine the context-aware vectors with the historical operation information to generate atomic operation instructions or atomic operation instruction sequences.

[0087] In the above step S137, by combining the weight information in the context-aware vectors with each historical operation instruction in the historical operation information, atomic operation instructions or atomic operation instruction sequences that can be directly executed can be generated.

[0088] In the above implementation process, starting from the input semantic vector sequence and historical operation information, through the processing of relative position encoding and the multi-head attention mechanism, the complex instruction processing channel can effectively capture the position information in the sequence and integrate it into the understanding and processing of complex instructions. This helps to more accurately parse instructions with complex syntactic structures and long-distance dependencies, and generate accurate operation instructions in combination with historical context. That is, it improves the processing speed and accuracy of speech, thereby improving the efficiency of voice control of the virtual platform.

[0089] Please continue to refer to Figure 4 , in some alternative embodiments, according to the complexity of the initialization instruction, generating an atomic operation instruction or a sequence of atomic operation instructions may further include: If it is determined that it has not exceeded, then execute step S138: The fast response channel performs a convolution operation based on the Mel-frequency cepstral coefficients of the initialization instruction to capture key acoustic patterns.

[0090] In the above step S138, first, MFCC (Mel-frequency cepstral coefficients) features can be extracted from the input initialization instruction. Since MFCC simulates the sensitivity of the human auditory system to different frequencies, MFCC can effectively represent the acoustic characteristics of speech signals. Specifically, it can be implemented as follows: First, convert the speech signal of the initialization instruction into a spectrum, then extract the energy through a Mel filter bank, and finally perform logarithmic transformation and discrete cosine transformation to obtain MFCC coefficients.

[0091] Then, in combination with Table 1, the MFCC features can be processed using convolutional layer 1. Specifically, it can be implemented as follows: The convolutional kernels in convolutional layer 1 slide on the MFCC feature sequence to perform convolutional operations. Among them, hyperparameters such as the size and stride of the convolutional kernel usually affect the dimension and details of the output features. For example, using a convolutional kernel with a size of 3×3 and a stride of 2 can reduce the dimension of the features while extracting local acoustic patterns.

[0092] Step S139: Perform global pooling dimensionality reduction on the initialization instruction after the convolution operation, and use an activation function to output the probability distribution of the elements in the atomic operation instruction set.

[0093] In the above step S139, first, the feature map after the convolution operation usually has a relatively high dimension, and a global pooling layer can be used to reduce the dimension of the feature map. Specifically, it can be implemented as follows: The global pooling layer performs operations such as averaging or taking the maximum value of all elements in each feature map to obtain a fixed-length vector. To reduce the dimension of the features, reduce the computational complexity, and at the same time retain key information.

[0094] Then, the dimension-reduced feature vectors are input into an activation function, such as: Sigmoid. The activation function maps the features to a specific numerical range. The activation function maps the output to between 0 and 1, which can represent a probability distribution and is used to convert the dimension-reduced features into the probability distribution of each element in a predefined atomic operation instruction set.

[0095] Step S1310: Generate an atomic operation instruction or an atomic operation instruction sequence according to the probability distribution.

[0096] In the above step S1310, first, since the output after being processed by the activation function is a probability distribution, which represents the possibility of each element in the predefined atomic operation instruction set corresponding to the input voice instruction being executed. For example, assume that the atomic operation instruction set includes operations such as "start titration", "record data", "add solution", etc. The output probability distribution will give the probability of each operation being executed. Then, according to the probability distribution, the atomic operation instruction with the highest probability is selected as the atomic operation instruction or the atomic operation instruction sequence and output.

[0097] In the above implementation process, through the convolution operation of the Mel-frequency cepstral coefficients on the initialization instruction, global pooling for dimension reduction, the output of the probability distribution by the activation function, and the output of the atomic operation instruction or the atomic operation instruction sequence in combination with the probability distribution, the fast processing of simple instructions is realized, and the efficient recognition and fast response of simple instructions are realized. Finally, the processing speed and accuracy of speech are also improved, thereby improving the efficiency of voice control of the virtual platform.

[0098] Please refer to Figure 5 , Figure 5 is the specific flowchart of step S140 in the voice control method of the virtual platform provided by the embodiment of the present application. In some alternative embodiments, step S140 may include: Step S141: Construct a four-dimensional state tensor for describing the virtual platform screen.

[0099] In the above step S141, the tensor formula is: ; In the formula, x, y, z represent the scene coordinate system of the virtual platform screen, which can represent the position information of each object in the virtual environment represented by the virtual platform, such as the position targets of experimental equipment, chemical reagent bottles, virtual operation tables, etc. This coordinate system can specifically be the WebGL scene coordinate system. t represents the time stamp, which can be used to record the time information of the experimental operation and can help track the progress of the experiment and the timing relationship of the operations. v represents the state of the operation object corresponding to the atomic operation instruction or the atomic operation instruction sequence, such as: the volume, temperature, concentration of the solution, the switch state of the equipment, etc. c pRepresents the operation stages divided for the operation process, such as: "preparation stage", "titration stage", "heating stage", etc. r represents the constraint rule, r ∈ {0, 1} 63 Activation state, such as max_temp = 100 °C.

[0100] Step S142: Extract the image features of the virtual platform screen to obtain an image feature vector.

[0101] In the above step S142, exemplarily, an experimental scene image generated by a WebGL renderer, such as: a 224×224×3 RGB image, is input into the ResNet-50 model to extract image features, and finally an image feature vector is generated.

[0102] Step S143: Perform a linear projection on the atomic operation instruction or atomic operation instruction sequence to obtain a projected speech feature vector.

[0103] In the above step S143, the linear projection can be to transform the original vector into a new vector space through matrix multiplication. In the embodiments of the present application, it can be to map the semantic vector of the speech instruction to a space matching the image feature vector. The obtained speech feature vector is usually in the same space as the image feature vector and can be meaningfully compared, fused, etc. with the image features.

[0104] Step S144: Combine the image feature vector and the projected speech feature vector to perform cross-modal attention calculation to obtain an attention weight distribution.

[0105] In the above step S144, the formula for cross-modal attention calculation is: ; In the formula, Q i is the linear projection of the speech feature vector, K j is the linear projection of the image feature vector, and d is the dimension of the speech feature vector and the image feature vector.

[0106] Attention calculation can be regarded as a process of query (Query), key (Key), and value (Value). Specifically, the speech feature vector can be used as the query, the image feature vector as the key and value, calculate the similarity score between the query and the key, and use these scores to perform weighted summation on the value, so as to determine the relevance degree of the operation mentioned in the speech instruction to each region in the image.

[0107] Step S145: Combine the attention weight distribution and the four-dimensional state tensor to verify the rationality of the atomic operation instruction or atomic operation instruction sequence.

[0108] In the above step S145, firstly, the region related to the instruction in the image can be found through the attention weight distribution, and combined with the scene coordinate system information in the four-dimensional state tensor, to confirm whether the spatial position of the operation object is reasonable. For example, if the instruction is "add solution", check whether the attention weight is concentrated on the position of the virtual beaker.

[0109] Secondly, the timestamp information in the four-dimensional state tensor can be used to verify whether the instruction conforms to the timing logic of the experiment. For example, it is reasonable to receive the "record data" instruction after the "heating stage".

[0110] Thirdly, according to the experimental object state information, verify whether the current state of the operation object is suitable for executing the instruction. For example, check whether the solution has reached a stable state that can be recorded.

[0111] Fourthly, combined with the experimental stage information, confirm whether the instruction matches the current stage. For example, the "centrifugation" instruction is reasonable in the "mixing stage", but may not be reasonable in the "observation stage".

[0112] In the above implementation process, through the attention weight distribution, the specific position or object in the virtual experimental platform referred to by the current voice instruction (atomic operation instruction or atomic operation instruction sequence) can be determined, so as to verify the rationality of the atomic operation instruction or atomic operation instruction sequence, ensuring the rationality, feasibility and safety of the instruction in the experimental environment. Thus, it further improves the processing speed and accuracy of speech, and then improves the efficiency of voice control of the virtual platform.

[0113] Please refer to Figure 6 , Figure 6 which is the specific flowchart of step S145 in the voice control method of the virtual platform provided by the embodiment of the present application. In some optional embodiments, step S145 may include: Step S1451: Combine the attention weight distribution and the four-dimensional state tensor to verify the atomic operation instruction or atomic operation instruction sequence according to natural laws.

[0114] Step S1452: Compare the object positioning heat map corresponding to the attention weight distribution with the four-dimensional state tensor to verify whether the operation object corresponding to the atomic operation instruction or atomic operation instruction sequence is consistent with the actual object in the virtual platform screen.

[0115] In the above step S1452, the relevant region in the object positioning heat map can be compared with the coordinate system in the four-dimensional state tensor to confirm whether the operation object in the instruction is consistent with the actual object in the virtual platform. For example, check whether the "add solution" instruction corresponds to the position of the beaker in the heat map.

[0116] In the above implementation process, by comparing the relevant regions in the object positioning heat map with the coordinate system in the four-dimensional state tensor, it is confirmed whether the operation object in the instruction is consistent with the actual object in the virtual platform, so as to verify the instruction. Further ensure the rationality, feasibility and safety of the instruction in the experimental environment, etc. Thereby further improving the processing speed and accuracy of the voice, and thus improving the efficiency of voice control of the virtual platform.

[0117] Please refer to Figure 7 , Figure 7 FIG. is a schematic diagram of a voice control system of a virtual platform provided by an embodiment of the present application. Based on the same concept, an embodiment of the present application provides a voice control system 700 for a virtual platform, including: A voice input layer 710, configured to perform noise reduction and start and end point recognition on the acquired initial voice instruction to obtain a processed voice instruction; A multi-language processing layer 720, configured to map the text converted from the processed voice instruction to a predefined atomic operation instruction set to obtain an initialization instruction; A reflective neural network 730, configured to generate an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction; wherein, the atomic operation instruction and the atomic operation instruction sequence are respectively instructions that can be directly executed; A spatio-temporal perception matrix 740, configured to verify the rationality of the atomic operation instruction or the atomic operation instruction sequence by combining the content currently virtualized by the virtual platform and the natural law; An execution layer 750, configured to execute the verified atomic operation instruction or atomic operation instruction sequence.

[0118] It should be understood that this system corresponds to the above-mentioned embodiment of the voice control method for a virtual platform and can execute each step involved in the above method embodiment. The specific functions of this system can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. This system may include at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the device.

[0119] An embodiment of the present application further provides a storage medium, which includes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the above method.

[0120] Among them, the computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.

[0121] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0122] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0123] The above description is only an optional implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the embodiments of the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application.

Claims

1. A voice control method for a virtual platform, characterized in that: include: Perform noise reduction and start and end point recognition on the acquired initial voice command to obtain a processed voice command; Mapping the text converted from the processed voice instruction to a predefined atomic operation instruction set to obtain an initialization instruction; Generate an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction; wherein the atomic operation instruction and the atomic operation instruction sequence are instructions that can be directly executed; Verify the rationality of the atomic operation instruction or atomic operation instruction sequence in combination with the content currently virtualized by the virtual platform and the laws of nature; and Execute the verified atomic operation instruction or atomic operation instruction sequence.

2. The method according to claim 1, characterized in that The step of performing noise reduction and start and end point recognition on the acquired initial voice command to obtain a processed voice command includes: Representing the initial voice command as a time domain signal; and The noise component in the initial voice command is removed by using spectral subtraction; wherein the implementation formula of the spectral subtraction is: ; In the formula, f represents frequency, X clean (f) represents the pure voice signal after the noise component is removed from the initial voice command, X(f) represents the frequency domain signal after the time domain signal is Fourier transformed, N(f) represents the preset room environment noise template; α is the over-reduction factor, β is the spectrum lower limit protection coefficient, Indicates phase information of the initial voice command.

3. The method according to claim 2, characterized in that The step of performing noise reduction and start and end point recognition on the acquired initial voice command to obtain a processed voice command further includes: Convert the frequency spectrum of the frequency domain signal after inverse Fourier transformation into a time domain signal; Determine whether the time domain signal satisfies and Whether the number of consecutive frames exceeds a first frame number threshold; If the determination is satisfied, the first frame in the continuous number of frames is used as the starting point of the voice command; Determine whether the time domain signal satisfies and Whether the number of consecutive frames exceeds a second frame number threshold; If the condition is satisfied, the last frame in the continuous number of frames is used as the end point of the voice command; Among them, E n With Z n They are calculated by the following formulas: ; ; In the formula, E n represents the short-time energy of the nth frame in the time domain signal, L represents the frame length, x(k) represents the amplitude of the kth sampling point, Z n represents the zero-crossing rate of the nth frame in the time domain signal, and I represents the judgment function, which is used to judge Is it true? μE and μZ represent the mean of the short-time energy and the mean of the zero-crossing rate of the noise component respectively; σE and σZ represent the standard deviation of the short-time energy and the standard deviation of the zero-crossing rate of the noise component.

4. The method according to claim 1, characterized in that: Mapping the text converted from the processed voice instruction to a predefined atomic operation instruction set to obtain an initialization instruction includes: The processed voice instructions are framed and input into a pre-built semantic feature extraction module to generate a phoneme sequence; Inputting the phoneme sequence into a pre-built semantic feature alignment module to generate a semantic vector; and The semantic vector is subjected to normalized mapping to obtain the initialization instruction; wherein the formula of the normalized mapping is: ; In the formula, z 标 Indicates the initialization instruction, h CLS is the semantic vector, z is a predefined atomic operation instruction set, and z i is an element in the atomic operation instruction set, E(z i ) represents the pre-trained instruction embedding vector.

5. The method according to claim 1, characterized in that Generating an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction includes: The instruction complexity of the initialization instruction is calculated by a reflective neural network; wherein the reflective neural network includes a fast response channel and a complex instruction processing channel; the calculation formula of the instruction complexity is as follows: ; In the formula, lexical_complexity represents the lexical complexity of the initialization instruction, syntactic_depth represents the syntactic depth of the initialization instruction, and Score represents the instruction complexity; Determining whether the complexity of the instruction exceeds a complexity threshold; If it is determined that the number exceeds the limit, the complex instruction processing channel receives the initialization instruction and the cached historical operation information, and performs embedding layer conversion on the initialization instruction to output an embedded vector; The relative position coding of the initialization instruction is calculated using a position coding formula; wherein the position coding formula is: ; Wherein, k represents the dimension index, d represents the embedding dimension, i and j represent the sequence bits of the two elements in the initialization instruction respectively; Fusion the embedding vector with the relative position code to obtain a fused embedding vector; Processing the fused embedding vector using a multi-head attention mechanism to obtain a context-aware vector; and The context-aware vector is combined with the historical operation information to generate the atomic operation instruction or atomic operation instruction sequence.

6. The method according to claim 5, characterized in that The step of generating an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction further includes: If it is determined that it does not exceed, the fast response channel performs a convolution operation based on the Mel-frequency cepstral coefficients of the initialization instruction to capture the key acoustic mode; Performing global pooling dimensionality reduction on the initialization instructions that have undergone the convolution operation, and outputting the probability distribution of elements in the atomic operation instruction set using an activation function; and The atomic operation instruction or atomic operation instruction sequence is generated according to the probability distribution.

7. The method according to claim 1, characterized in that The checking of the rationality of the atomic operation instruction or atomic operation instruction sequence in combination with the current virtualized content of the virtual platform and the laws of nature includes: Construct a four-dimensional state tensor for describing the virtual platform screen; wherein the tensor formula is: ; Wherein, x, y, z represent the scene coordinate system of the virtual platform screen, t represents the timestamp, v represents the state of the operation object corresponding to the atomic operation instruction or atomic operation instruction sequence, and c p It represents the operation phases of the operation process; r represents the constraint rules; Extracting image features of the virtual platform screen to obtain an image feature vector; Performing linear projection on the atomic operation instruction or the atomic operation instruction sequence to obtain a projected speech feature vector; Combining the image feature vector with the projected speech feature vector, a cross-modal attention calculation is performed to obtain an attention weight distribution; wherein the formula for cross-modal attention calculation is: ; In the formula, Q i is the linear projection of the speech feature vector, K j is the linear projection of the image feature vector, and d is the dimension of the speech feature vector and the image feature vector; In combination with the attention weight distribution and the four-dimensional state tensor, the rationality of the atomic operation instruction or the atomic operation instruction sequence is verified.

8. The method according to claim 7, characterized in that The checking the rationality of the atomic operation instruction or the atomic operation instruction sequence in combination with the attention weight distribution and the four-dimensional state tensor includes: combining the attention weight distribution and the four-dimensional state tensor, and verifying the atomic operation instruction or the atomic operation instruction sequence according to the natural law; and By comparing the object positioning heat map corresponding to the attention weight distribution with the four-dimensional state tensor, it is verified whether the operation object corresponding to the atomic operation instruction or the atomic operation instruction sequence is consistent with the actual object in the virtual platform screen.

9. A voice control system for a virtual platform, characterized in that: include: The voice input layer is used to reduce noise and identify the start and end points of the acquired initial voice commands to obtain processed voice commands; A multi-language processing layer, used for mapping the text converted from the processed voice instruction to a predefined atomic operation instruction set to obtain an initialization instruction; A reflective neural network, used to generate an atomic operation instruction or an atomic operation instruction sequence according to the complexity of the initialization instruction; wherein the atomic operation instruction and the atomic operation instruction sequence are instructions that can be directly executed; The time-space perception matrix is ​​used to verify the rationality of the atomic operation instruction or atomic operation instruction sequence in combination with the content currently virtualized by the virtual platform and the laws of nature; The execution layer is used to execute the verified atomic operation instruction or atomic operation instruction sequence.

10. A storage medium, characterized in that: The storage medium comprises a computer-readable storage medium; a computer program is stored on the computer-readable storage medium, and the computer program executes the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Signal acquisition and processing system with speech recognition

    CN108735206A

  • User online questioning processing method and system based on complexity analysis

    CN111931498A

  • Voice control method and system for automatically broadcasting contents on large screen

    CN112102828A

  • PLC speech recognition method based on machine learning

    CN113643692A

  • Command word response method, control equipment and device

    CN115731923A