Voice interaction method and system based on 3D virtualization
By combining user historical commands and 3D environment context speech recognition model to generate a set of candidate prediction instructions, the existing system's inaccurate command generation and slow response speed in complex environments is solved, and higher speech interaction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510857334.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing voice recognition systems cannot effectively combine user historical commands and 3D environment context in complex environments, resulting in inaccurate command generation and slow response speed.
By receiving the user's real-time voice for preprocessing, a speech recognition model is established, the user's pronunciation characteristics and clarity are analyzed, and the user's historical commands and the current 3D environment context is combined. The optimized speech recognition model is used to generate a set of candidate prediction instructions, and the priority of each candidate instruction is identified to quickly respond to user needs.
It significantly improves the accuracy and real-time response capabilities of voice interactions, can effectively deal with problems such as unclear pronunciation or voice interference, and improves the robustness and adaptability of the system.
Smart Images

Figure CN120388565A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice interaction technology, and specifically to a 3D virtual-based voice interaction method and system. Background Art
[0002] With the rapid development of computer technology and artificial intelligence, speech recognition technology has become a key means of human-computer interaction. Especially in immersive environments such as virtual reality (VR) and augmented reality (AR), speech recognition not only provides users with a more natural way of interaction but also significantly enhances the user experience. However, existing speech recognition systems still face many challenges in real-world applications, such as varying speech clarity, background noise interference, adaptability issues for users with articulation disorders, and understanding context in complex environments.
[0003] Traditional speech recognition systems are typically based on static models, utilizing fixed features and algorithms to process speech signals. While these systems can handle standard speech input, they often struggle with variations in pronunciation, accent, speech rate, and ambient noise. Furthermore, existing speech recognition models lack sufficient adaptability to handle complex scenarios and contextual changes. For example, information about a user's behavior, context, and historical commands in a virtual environment is often not effectively utilized, limiting the speed of command generation and response, impacting the user's interactive experience.
[0004] While some current optimization methods have introduced technologies like deep learning to improve speech recognition accuracy, most are limited to analyzing single speech input and neglect the combined use of historical user commands and environmental context. Specifically, existing technologies have yet to effectively combine historical user interactions with real-time environmental context to dynamically optimize speech recognition models and the generation of candidate command sets, further improving the accuracy and responsiveness of speech recognition systems in practical applications. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by the present invention is that the existing speech recognition system cannot effectively combine the user's historical commands and 3D environment context in a complex environment, resulting in inaccurate command generation and slow response speed.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: a 3D virtual-based voice interaction method, comprising: receiving and preprocessing a user's real-time voice, establishing a voice recognition model, and analyzing the user's pronunciation characteristics and clarity;
[0008] Optimize the speech recognition model based on the user's pronunciation features and clarity;
[0009] Combine the user's historical commands and the current 3D environment context, and use the optimized speech recognition model to generate a set of candidate prediction instructions for the user's next step; the set of candidate prediction instructions includes multiple candidate instructions;
[0010] Identify the priority of each candidate instruction, and use the candidate instruction with the highest priority as the predicted instruction for the user's next step. Preferentially load the relevant resources of the predicted instruction to quickly respond to the user's needs; the identification of the priority of each candidate instruction includes calculating the priority of each candidate instruction according to the speech clarity score, the execution frequency of historical commands, and the matching degree of the instruction with the current environment.
[0011] As a preferred solution of the 3D virtual-based speech interaction method described in the present invention, wherein: the preprocessing includes using the Wiener filtering algorithm to remove the noise of the real-time speech, and dividing the denoised real-time speech into multiple consecutive frame audios, and the overlapping part between each frame is ; use the Mel-frequency cepstral coefficients to extract the time-frequency features of each frame of audio and establish a time-frequency feature matrix; wherein, represents the overlapping length;
[0012] The time-frequency features include the spectrum, pitch, volume, fundamental frequency, and spectral flatness of the audio.
[0013] As a preferred solution of the 3D virtual-based speech interaction method described in the present invention, wherein: the speech recognition model includes a convolutional neural network layer, a pooling layer, and a fully connected layer;
[0014] Input the time-frequency feature matrix into the convolutional neural network layer for convolution operation to extract local features and form a local feature map; the pooling layer uses the maximum pooling method to reduce the dimension of the local features and obtain a key local feature map; the fully connected layer converts the local feature map and the key local feature map into feature vectors and constructs them into a global feature map.
[0015] As a preferred solution of the 3D virtual-based speech interaction method described in the present invention, wherein: analyzing the pronunciation clarity of the user includes using the wavelet transform method to extract the high-frequency coefficients and low-frequency coefficients from the global feature map; calculating the energy of the high-frequency coefficients and the energy of the low-frequency coefficients; based on the ratio of the high-frequency energy to the low-frequency energy, determine the clarity score of the speech; set a clarity score threshold, when the clarity score is less than the clarity score threshold, it is determined that the speech clarity is insufficient, and the error tolerance rate is increased;
[0016] Analyzing the user's pronunciation features includes using a sliding window method to calculate the local energy of high-frequency detail coefficients, finding the local energy peak, and determining the stressed part; calculating the average time interval between adjacent local energy peaks; setting a fast speech rate threshold and a slow speech rate threshold. When the average time interval is less than the fast speech rate threshold, it is determined that the speech rate is too fast, and the weight of the high-frequency filter in the convolutional neural network is increased to enhance the ability to capture rapidly changing syllables; when the average time interval is greater than the slow speech rate threshold, it is determined that the speech rate is too slow, and the weight of the low-frequency filter in the convolutional neural network is increased to enhance the recognition ability of slow-paced speech.
[0017] As a preferred embodiment of the 3D virtual-based voice interaction method of the present invention, wherein: the generation of the candidate prediction instruction set for the user's next step includes receiving the user's historical instruction sequence; using a long short-term memory network to convert the historical command sequence into temporal features, generating the hidden state of the historical command, and recursively updating the hidden state to learn the information in the historical instruction to obtain the historical command state;
[0018] Using the optimized speech recognition model, analyzing the received real-time speech input to obtain real-time speech features; constructing a multi-layer deep model to deduce the possible operation intentions of the user layer by layer and generate a candidate prediction instruction set for the user's next step.
[0019] As a preferred embodiment of the 3D virtual-based voice interaction method of the present invention, wherein: the multi-layer deep model includes an LSTM layer, a GCN layer, and an output layer;
[0020] In the LSTM layer, combining the historical command state and the real-time speech features, using LSTM to generate a preliminary candidate instruction set; combining the historical command state with the real-time speech features, using LSTM to generate a real-time candidate instruction set;
[0021] The GCN layer includes pre-training a rule engine offline using a meta-learning framework based on the user's historical instruction sequence to establish a user-specific rule model; when the user starts voice interaction, based on the current instruction speech clarity score and the average value of all syllable intervals in the current instruction, sampling a subsample from the historical interaction data with the same instruction speech clarity score and the average value of all syllable intervals in the current instruction to construct a support set, and quickly fine-tuning the meta-model parameters to generate a rule engine personalized for the current user; for each candidate instruction, extracting a semantic-behavior joint feature vector using the support set; based on the semantic-behavior joint feature vector, constructing a heterogeneous graph spectrum; the node set of the constructed heterogeneous graph spectrum includes: instruction nodes, target object nodes, and context semantic nodes; for any two nodes i, j in the heterogeneous graph spectrum, the corresponding semantic-behavior joint feature vectors are and When , an edge is established between i and j, and the weight of the edge is ; where represents the feature similarity, represents the set structural connection threshold.
[0022] As a preferred solution of the 3D virtual-based voice interaction method described in the present invention, wherein: the heterogeneous graph spectrum is subjected to structural propagation through a multi-layer graph convolutional network to obtain a graph structure embedding representation of the candidate instruction, and the instruction prediction result is output through a prediction head structure; the prediction result is compared with the true interaction behavior label to identify the prediction error. When the prediction error exceeds the set threshold, the structural risk factor of the candidate instruction is extracted to construct a structure attribution matrix;
[0023] The structural risk factor includes a semantic ambiguity factor and a path conflict factor; the semantic ambiguity factor is a quantization index of the semantic concentration degree when the candidate instruction node points to multiple objects or semantic nodes; all edges of the candidate instruction node and their corresponding edge weights are extracted, and the edge weight set is normalized to a probability distribution, and then the uncertainty degree of the probability distribution is calculated based on the Shannon entropy, and the obtained entropy value is used as the value of the semantic ambiguity factor after normalization processing; the path conflict factor represents the overlapping degree of the path where the candidate instruction node is located and the adjacent candidate path in the graph structure on the structural edge set. Define the path where the node is , extract 's edge set ; construct a neighbor path set according to the instruction nodes directly connected to in the graph, and extract the edge set of each neighbor path respectively, and calculate the structural Jaccard similarity between , and take the maximum value in the structural Jaccard similarity as ;
[0024] Taking the structure attribution matrix as a regulatory factor, introduce it into the parameter update path of the personalized rule engine, and use the updated personalized rule engine to calculate the personalized preference scores of each candidate instruction, guiding the meta-model to avoid high-structure-risk areas in subsequent tasks; fuse the personalized preference scores and semantic adaptation scores to generate the comprehensive scores of the candidate instructions, sort them according to the scores from high to low, and select the top M candidate instructions as the environmental context candidate instruction set; the semantic adaptation score is the embedding probability value output by the prediction head after the graph convolutional network structure propagates, and is used to measure the structural semantic fitting degree between the candidate instruction and the target object and context nodes in the graph; the output layer uses the attention mechanism to perform weighted fusion on the preliminary candidate instruction set, the real-time candidate instruction set, and the environmental context candidate instruction set to generate the candidate prediction instruction set for the user's next step.
[0025] A voice interaction system based on 3D virtuality, wherein:
[0026] A data module that receives the user's real-time voice, preprocesses it, establishes a speech recognition model, and analyzes the user's pronunciation features and clarity;
[0027] An optimization module that optimizes the speech recognition model based on the user's pronunciation features and clarity;
[0028] A prediction module that combines the user's historical commands and the current 3D environmental context, and uses the optimized speech recognition model to generate a candidate prediction instruction set for the user's next step; the candidate prediction instruction set includes multiple candidate instructions;
[0029] A loading module that identifies the priority of each candidate instruction, takes the candidate instruction with the highest priority as the prediction instruction for the user's next step, preferentially loads the relevant resources of the prediction instruction, and quickly responds to the user's needs.
[0030] A computer device, comprising: a memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method described in any one of the present inventions.
[0031] A computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, it implements the steps of the method described in any one of the present inventions.
[0032] Advantages of the present invention: The voice interaction method based on 3D virtual provided by the present invention proposes a candidate instruction generation method for a multi-layer deep model by combining an optimized voice recognition model, user historical commands, and 3D environment context, significantly improving the accuracy and real-time response ability of voice interaction. Through a multi-layer deep learning structure to gradually reason and fuse voice input, historical data, and environmental information, the present invention can dynamically optimize the instruction generation process to ensure that the system can more accurately predict user needs in complex scenarios. In addition, by adopting a fault tolerance rate adjustment mechanism and an optimization strategy based on clarity, the system can effectively cope with problems such as unclear pronunciation or voice interference, improving the robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 It is the overall flowchart of a voice interaction method based on 3D virtual provided by the first embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0036] Embodiment 1, referring to Figure 1 , which is an embodiment of the present invention, provides a voice interaction method based on 3D virtual, including:
[0037] S1: Receive the real-time voice of the user and perform preprocessing, establish a voice recognition model, and analyze the pronunciation characteristics and clarity of the user; based on the pronunciation characteristics and clarity of the user, optimize the voice recognition model.
[0038] Use the Wiener filtering algorithm to remove the noise of the received real-time voice, and then divide the denoised real-time voice into multiple frames with a duration of , and the overlapping part between each frame is ; where represents the frame length, represents the overlapping length, and is Half of it. Ensure that the signal within each frame can be regarded as a stationary signal, which is convenient for time-frequency feature extraction. After the time-frequency transformation of each frame of audio signal, relevant time-frequency features are extracted to provide accurate input for the subsequent speech recognition process.
[0039] The speech recognition model includes a convolutional neural network layer, a pooling layer, and a fully connected layer.
[0040] Input the time-frequency feature matrix into the convolutional neural network layer, and perform a convolution operation on the time-frequency feature matrix through a convolution kernel to extract local features. The formula is:
[0041] After the convolution operation, use max pooling in the pooling layer to reduce the dimension of the local features and retain the most important feature information to obtain the key local feature map. The formula is expressed as:
[0042]
[0043] Among them, represents the eigenvalue of the local feature map output by the convolutional layer at position ; represents the element at position in the input time-frequency feature matrix ; represents the weight at position in the convolution kernel ; is the convolution kernel; represents the index of the convolution kernel. represents the eigenvalue of the key local feature map output by the pooling layer at position ; represents the eigenvalue at position in the local feature map output by the convolutional layer, is the local feature map after being processed by the convolutional layer; represents the offset of the pooling window. and represent the position index.
[0044] The fully connected layer transforms the local feature map and the key local feature map into feature vectors to construct global features. The fully connected layer connects multiple neurons to all other neurons and maps the high-dimensional feature information to the final output space.
[0045]
[0046] Flatten and into a feature vector , which is used as the input of the fully connected layer.
[0047]
[0048]
[0049] Among them, v represents the flattened feature vector; represents the local feature map output by the convolutional layer; represents the key local feature map output by the pooling layer; Flatten represents the flattening operation; z represents the linear output of the fully connected layer; represents the weight of the fully connected layer, represents the bias term; represents the global feature map, represents the activation function.
[0050] Perform on the global feature map layer discrete wavelet transform to obtain the low-frequency coefficients and high-frequency coefficients of the layer, and calculate the energy of the high-frequency detail part and the energy of the low-frequency approximation part of the layer respectively:
[0051]
[0052]
[0053] Calculate the clarity score based on the high-frequency energy and the low-frequency energy:
[0054]
[0055] Among them, is a very small positive number to prevent the denominator from being zero. Set the clarity score threshold When it is considered that the speech clarity is insufficient, and the fault tolerance rate adjustment mechanism of the recognition model is started.
[0056] Among them, represents the low-frequency approximation coefficient after the layer wavelet transform, which contains the stationary components of the speech signal and usually represents the basic contour or main energy of the speech; represents the high-frequency detail coefficient after the layer wavelet transform, which contains the changing parts, rapid changes or detailed information of the speech, such as stress, intonation changes, pronunciation mutations, etc.; represents the global feature map; represents the operation of performing the layer discrete wavelet transform on the input signal ; represents the Energy of the high-frequency detail part of the layer; Represents the high-frequency coefficient after wavelet transform of the layer, and the th element; Represents the energy of the low-frequency approximation part of the layer; Represents the low-frequency coefficient after wavelet transform of the layer, and the th element;
[0057] Represents the speech intelligibility score. Set a sliding window on the high-frequency detail coefficient , and calculate the local high-frequency energy within the th window:
[0058]
[0059] Use the sliding comparison method to detect the energy peak. When the following conditions are met: , and , it is marked as a valid peak. Extract all valid peaks to obtain the peak time point sequence , and calculate the adjacent syllable time intervals based on this:
[0060]
[0061] Further calculate the average value of all syllable intervals:
[0062]
[0063] Set the fast speech rate threshold and the slow speech rate threshold . Judge the speech rate according to the following rules: If , it is determined that the speech rate is too fast; if , it is determined that the speech rate is too slow; if , it is determined that the speech rate is normal. When the speech rate is too fast or too slow, trigger the weight optimization mechanism. Represents the number of syllables detected in this voice segment, the time interval number index; Represents the th syllable and the time interval between it and the next syllable.
[0064] Among them, Represents the local high-frequency energy within the th sliding window ; Represents the th sliding window; Indicates the high-frequency detail coefficient at position after the -layer wavelet transform; Indicates the local high-frequency energy of the -th sliding window; Indicates the local high-frequency energy of the -th sliding window; Indicates the energy threshold; Indicates the time point corresponding to the -th effective peak; Indicates the time point corresponding to the -th effective peak; Indicates the time interval between the -th syllable and the -th syllable. Indicates the number of effective peaks; Indicates the average value of all syllable intervals; Indicates the threshold for fast speech rate. Indicates the threshold for slow speech rate.
[0065] During the training process, the backpropagation algorithm is used to calculate the gradient of the loss function, and the network parameters are optimized through gradient descent. The cross-entropy loss function is used to measure the gap between the model output and the actual labels, and the recognition accuracy is improved by optimizing the loss function.
[0066] The fault tolerance rate adjustment mechanism includes that for low-clarity speech, by increasing the convolutional kernel size, the model can capture more local features when processing low-clarity speech,
[0067] where is the original convolutional kernel, is the magnification factor, is the new convolutional kernel.
[0068] And the stride of the pooling layer is adjusted to 1 / 2 of the original stride, so that the pooling window can capture more fine-grained features and improve the recognition ability for low-clarity speech.
[0069]
[0070] where is the new pooling stride, is the original stride.
[0071] The weight optimization mechanism includes that through the backpropagation algorithm, more attention is paid to the high-error regions during training, especially the feature regions in fast or slow speech. The weights of the convolutional kernel are updated using the gradient descent method to make the model more sensitive to these regions.
[0072]
[0073] Among them, is the updated convolutional kernel weight, is the current convolutional kernel weight, is the learning rate, is the gradient of the loss function with respect to the weight.
[0074] According to the speed of speech, adjust the window size of the pooling layer. When the speech speed is too fast, increase the pooling window to ensure that the features in fast speech can be effectively extracted. When the speech speed is too slow, decrease the pooling window so that the detailed features of slow speech can be more accurately extracted. The formula for adjusting the window size of the pooling layer is expressed as:
[0075]
[0076] Among them, is the adjusted pooling window size, is the original pooling window size.
[0077] By introducing a multi-layer deep learning model and an optimization mechanism, the robustness and accuracy of the speech recognition system have been significantly improved. Compared with traditional speech recognition systems, our improvements are mainly reflected in: on the one hand, by combining convolutional neural networks (CNNs) with time-frequency feature extraction, fine-grained analysis of the input speech is carried out to effectively extract local features of the audio, ensuring that the system can accurately recognize speech with unclear and indistinct pronunciation; on the other hand, by introducing multi-layer wavelet transforms, the speech clarity and pronunciation features can be dynamically evaluated, further adjusting the error tolerance rate of the system to cope with complex speech inputs. This innovation enables the system to maintain high-precision recognition ability in the face of non-standard pronunciation, speech impairments, or noise interference.
[0078] Furthermore, speech preprocessing and recognition optimization have significant practical value in 3D speech interaction systems. Background noise in real-time speech is removed by Wiener filtering, and then the processed speech is divided into overlapping frames to ensure local stationarity of the signals within the frames, which helps to extract more representative time-frequency features. The convolutional neural network structure extracts and compresses local features on this basis, obtains key speech patterns through pooling operations, and at the same time the fully connected layer further embeds the local maps into the global semantic space to form a unified feature vector. The speech clarity is measured by the energy ratio of the high-frequency and low-frequency coefficients decomposed by discrete wavelet transforms. When the score is lower than the set threshold, the adjustment mechanism of the recognition model structure is automatically triggered. This mechanism realizes the detailed capture of fuzzy speech information by expanding the receptive field of the convolutional kernel and compressing the pooling stride, significantly enhancing the adaptability of the model to non-standard pronunciation and low signal-to-noise ratio speech, and effectively improving the error tolerance and robustness of speech recognition.
[0079] In terms of speech rate feature recognition, a sliding window is used to analyze the high-frequency energy change, extract the syllable boundaries corresponding to the local peaks, and further calculate the time interval between adjacent peaks to evaluate the speech rate state. According to the comparison between the average time interval and the preset threshold, the system dynamically determines whether the speech rate is too fast or too slow. When the speech rate is too fast, the pooling window is enlarged to ensure the effective extraction of high-speed changing features; when the speech rate is too slow, the pooling window is reduced to capture subtle speech details. The recognition model uses the backpropagation algorithm during training, taking the recognition error regions generated at fast and slow speech rates as the focus of attention, and dynamically adjusting the weights of the convolutional kernels. This optimization strategy improves the model's response ability to speech rate fluctuations, enabling it to maintain a high recognition accuracy under different speech rate conditions. The overall structure has an adaptive rhythm adjustment function, which can effectively reduce the misrecognition rate caused by changes in pronunciation rhythm, and demonstrates higher stability and adaptability performance in multi-round interaction tasks.
[0080] S2: Combine the user's historical commands and the current 3D environment context, and use the optimized speech recognition model to generate a set of candidate prediction instructions for the user's next step; the set of candidate prediction instructions includes multiple candidate instructions.
[0081] Extract the user's historical instruction sequence from the user's interaction history, including the timestamp, execution frequency, and execution context information of the instructions.
[0082] Use LSTM to convert the historical command sequence into temporal features, which is expressed by the formula:
[0083] where, represents the hidden state at the th moment, represents the historical instruction input at the th moment. LSTM represents the recursive update of the hidden state to learn and store the information in the historical instructions. represents the hidden state at the th moment.
[0084] Use the optimized speech recognition model to analyze the real-time speech input to obtain real-time speech features .
[0085] The multi-layer deep model includes an LSTM layer, a GCN layer, and an output layer.
[0086] In the LSTM layer, the system analyzes the relationship between the historical command state and the current speech input, and combines the speech input to generate a preliminary set of candidate instructions:
[0087]
[0088] where, Represents the historical command state, which is the final state of the LSTM after learning the historical command sequence; Represents the real-time speech features, Represents the preliminary candidate instruction set.
[0089] Combine the historical command state with the real-time speech features, and use the LSTM to generate the real-time candidate instruction set:
[0090]
[0091] The GCN layer includes: obtaining the real-time 3D environmental context information, including the user's location, virtual target, interaction object, and scene state; providing the background of the current environment for the system to help the system make more intelligent predictions.
[0092] Offline, take the user's historical instruction sequence as a learning task for each user Conduct meta-training:
[0093]
[0094] Learning task Internally divided into two parts, including the support set Used to quickly adapt to the user's rule preferences; query set Used to evaluate whether the adapted model generalizes well.
[0095] During the support set training process of each meta-task: Encode the historical instruction sequence to generate a temporal state representation representing the user's behavior path; fuse the pronunciation clarity score, speech rate rhythm vector, and language feature embedding corresponding to each instruction; at the same time, embed the environmental state (location, object, scene state) where each instruction is located as a context vector; map the above multi-source data to the same feature space through a shared embedding network and perform joint optimization to obtain a semantic-behavior joint feature vector , expressed by the formula:
[0096]
[0097]
[0098] Among them, Represents the meta-model embedding function; Represents the current 3D environmental context vector; ups represents the coordinate or spatial position label of the user in the virtual scene; Represents the encoding information of the object associated with the current interaction instruction (such as pumping station, button, valve, etc.); Represents the overall state of the scene, such as spatial label, task stage, anomaly mark, etc.
[0099] In the MAML framework, the optimization objective of meta-training is set as follows: by simultaneously learning across multiple user tasks, the model is enabled to have the ability to generate a personalized rule engine exclusive to the current user quickly after receiving a small amount of individual user data so as to achieve efficient generalization and rapid adaptation. Among them represents the latest small sample training set of the user in the online stage; includes the speech clarity score, semantic label, and context object of each instruction; is the target instruction actually selected by the user recorded by the system; the data time window is limited to the last 5 - 10 sessions to ensure adaptability and real-time performance; the total number of data N is generally set to 8 - 16.
[0100] Inner loop: For each user task , perform gradient update on the support set to obtain the personalized parameters for this user:
[0101]
[0102] Among them, represents the current global model parameters; represents the inner loop learning rate; represents the loss function on the support set. represents the personalized model parameters after fine-tuning for user ; represents the gradient operation on the parameter ; represents the rule engine model with as the parameter.
[0103] Outer loop: Apply the adapted model of each task to the query set , calculate the generalization error, and update the meta-model parameters by aggregating the errors of all tasks and backpropagating:
[0104]
[0105] Among them, represents the meta-update learning rate; represents the value of the loss function on the query set; is the historical support set of the user is the historical query set of the user. : The gradient of the loss function with respect to the parameter ; represents the rule engine model with as the parameter.
[0106] When the user starts voice interaction, the system reads the interaction data of several recent rounds and constructs a small-sample training set. . Among them, instead of directly using all historical data for the training set, a feature guidance mechanism based on the current voice state and environmental context is introduced. From the historical interaction records, a subsample with the same average of speech clarity scores and all syllable intervals in the instruction as the current round of interaction is sampled as the support set to perform a quick fine-tuning on the meta-model parameters to obtain a personalized rule engine for the current user for real-time inference:
[0107]
[0108] Among them, represents the user 's exclusive rule engine. represents the parameter update path of the user personalization fine-tuning process in the meta-learning framework; starting from the global meta-model parameters , first calculate the loss function of the model on the support set of the user , and then take the gradient of this loss function with respect to the parameter to obtain the direction that can optimize the model performance best under the current data.
[0109] Construct a heterogeneous graph spectrum containing instruction entities, scene objects, and semantic labels ; the nodes include: instruction nodes, target object nodes, and context semantic nodes. Among them, represents the node set; represents the edge set.
[0110] The instruction node represents the semantic-behavior joint feature vector of each candidate voice instruction; the target object node represents the interaction entity in the current 3D scene; the context semantic node represents the context clues in the user's recent interaction, such as the current scene state, user location label, and previous operation path.
[0111] The connections between nodes are not predefined fixed edge types, but the edge relationships are dynamically established by calculating the similarity between their semantic-behavior feature vectors; for any two nodes i, j, their corresponding feature vectors are and , then the condition for establishing an edge between the two nodes is:
[0112]
[0113] Among them, is the improved feature similarity, Represents the structural connection threshold set by the system.
[0114] Edges include: instruction–object edge: indicates that the action intention contained in the candidate instruction is strongly correlated with the interaction of a certain object; instruction–instruction edge: indicates that there is high behavioral similarity or contextual semantic commonality between two candidate instructions (such as sequential operations or synonymous instructions); instruction–context edge: indicates that the current instruction has strong adaptability with the user's recent behavior context in the meta-learning feature space; object–context edge: indicates that a certain object is only frequently operated or concerned in a specific scenario state.
[0115] Adopt a multi-layer GCN propagation structure for context semantics, and transform the graph structure into an initial vector in the embedding space :
[0116]
[0117] Obtain candidate instructions The graph structure embedding of , Indicates the last layer of the GCN.
[0118] Among them, Represents the initial embedding vector of node in the graph, Represents node at the layer embedding vector; Represents the set of neighbor nodes of node ; Represents the weight of the layer GCN. Represents the normalization coefficient, Represents the activation function. Represents the embedding vector of the instruction node after the L-layer GCN.
[0119] Subsequently, introduce a prediction head structure, input the embedding vector into the fully connected layer, output the predicted instruction label , and then compare it with the true instruction label , calculate the prediction error , set a significant prediction error threshold according to historical experience, when is greater than the significant prediction error threshold, it is determined as an error instruction node.
[0120] Extract the structural risk factors of each candidate instruction to construct a structural attribution matrix; the structural risk factors include: the semantic ambiguity factor is that there are multiple equal-weight paths or discrete pointing node distributions in the graph structure for the candidate instruction, resulting in a semantic concentration degree lower than the set threshold; the path conflict factor: the candidate instruction path has a competitive relationship with other candidate paths, indicating a selection conflict risk.
[0121] The semantic ambiguity factor represents the quantification result of the semantic concentration degree when the candidate instruction node points to multiple objects or semantic nodes; obtain all the edge sets of the node , and extract the weight of each edge , normalize the edge set into a probability distribution , and then perform uncertainty calculation using Shannon entropy to obtain ; the factor is a continuous variable, and its value range is [0,1].
[0122] The path conflict factor represents the overlapping degree of the path where the candidate instruction node is located and the adjacent candidate path in the graph structure; define the path where the node is located as , extract the edge set of ; then construct a neighbor path set according to the instruction nodes directly connected to in the graph, and extract the edge set of each neighbor path , and calculate the structural Jaccard similarity between and .
[0123] The structural feature vector of each candidate instruction node is composed of the measurable semantic ambiguity factor and the path conflict factor in the graph structure, and is defined as:
[0124]
[0125] Multiply the attribution vector of each node by its error to form an attribution matrix , and the formula is expressed as:
[0126] Among them, represents the index of the candidate instruction node, represents the set of error instruction nodes.
[0127] The structure attribution matrix is not directly used for loss calculation, but to introduce the parameter update path of the meta-learning model. In the outer loop, the system uses to perform directional adjustment on the gradient of
[0128]
[0129] Among them, amplifies the gradient adjustment weight of the structure risk area, weakens the update direction of high-ambiguity and conflict paths, and forms an instruction embedding space with stronger structural stability and better generalization ability.
[0130] Using the parameter The updated user-specific rule engine calculates the preference score for each instruction:
[0131]
[0132] Among them, the output is a continuous value in the range of [0, 1], representing the probability that the user is inclined to execute the instruction in the current context; the higher the score, the more the instruction conforms to the user's behavior preference in similar historical contexts. represents the candidate instruction preference score.
[0133] Using an MLP (Multi-Layer Perceptron) to process the environmental state vector to obtain the embedding processing the current environmental context state:
[0134]
[0135] to obtain the graph structure semantic adaptation score of each candidate instruction in the current context:
[0136]
[0137] For each candidate instruction, fuse the meta-learning rule score and the graph structure semantic adaptation score:
[0138]
[0139] Among them, , which can be dynamically adjusted. represents the candidate instruction in the context adaptation score in the current graph structure. represents the final score of the candidate instruction after fusion.
[0140] Sort the final scores in descending order and select the top M output environmental context candidate instruction sets:
[0141]
[0142] The output layer includes: Through the attention mechanism, the information of historical commands, speech inputs, and environmental context is finally fused to generate an accurate set of candidate instructions.
[0143]
[0144] Among them, represents the set of candidate instructions. represents selecting the top M instructions in descending order of scores; represents the environmental context candidate instruction set.
[0145] In the candidate instruction generation stage, a multi-layer neural inference mechanism that combines meta-learning feature fine-tuning and graph-structured context modeling is innovatively constructed, significantly improving the response ability of the speech interaction system to personalized behavior patterns and complex 3D environmental changes. On the one hand, through the historical sample guidance mechanism based on the current speech state and syllable interval features, precise fine-tuning of the meta-model parameters is achieved, enabling the system to quickly generate a personalized rule engine that conforms to the current user preferences; on the other hand, by combining the structural semantic propagation in the heterogeneous graph spectrum and the graph neural network embedding representation, a structure-aware path is established between the instruction nodes, target objects, and context information, effectively eliminating the prediction deviation problem caused by the lack of context understanding in traditional static prediction models in multi-semantic and multi-path scenarios. This dual-path fusion mechanism not only improves the generation accuracy of candidate instructions but also maintains the generalization stability of the model under conditions of sudden user behavior changes or environmental switches.
[0146] Furthermore, by introducing a structural attribution mechanism for GCN prediction errors, by constructing a structural feature vector of candidate instructions (including semantic ambiguity factor and path conflict factor) and generating an attribution matrix, dynamic adjustment of the gradient direction during the model training process is achieved. This mechanism combines the node prediction errors in the graph with their structural risk factors to avoid interference from structurally ambiguous regions to the model learning path to the greatest extent, thus forming a closed-loop optimization feedback process of "error recognition - factor extraction - direction regulation". Compared with existing methods, this mechanism not only effectively alleviates the coupling interference problem between complex instruction paths but also improves the robustness and interpretability of the model in regions with low semantic concentration and high path overlap in the complex graph structure during the task adaptation stage, ultimately enhancing the actual deployment performance and human-machine collaborative intelligence level of the speech interaction system in high-dynamic virtual environments.
[0147] Existing methods generally rely on fixed templates or simplified context processing mechanisms, lacking the ability to adapt to the real-time pronunciation features and speech rate rhythms of users. This leads to problems such as response lags, command recognition deviations, and high repetition rates of candidate prediction commands in dynamic 3D virtual scenarios, and cannot effectively take into account structural factors such as semantic ambiguity and path conflicts.
[0148] By introducing a historical sample feature sampling mechanism that combines speech clarity and rhythm features, and combining with the meta-learning outer loop tuning process based on MAML, a highly personalized rule inference engine is constructed at the initial stage of task startup. At the same time, a structure attribution matrix is constructed based on the graph structure risk factors of candidate commands to guide the model to avoid high-risk paths during the structure embedding stage, significantly improving the generalization prediction ability in complex command networks. The experimental results show that on the 3D interaction simulation platform, it is superior to the existing model in terms of three indicators: semantic accuracy rate, command response speed, and structural stability, reflecting its system adaptability and practical application value in multi-round speech interaction tasks.
[0149] S3: Identify the priority of each candidate command, and use the candidate command with the highest priority as the user's next predicted command. Prioritize loading the relevant resources of the predicted command to quickly respond to the user's needs.
[0150] Calculate the priority of each candidate command according to the speech clarity score, the execution frequency of historical commands, and the matching degree between the command and the current environment. The formula is expressed as:
[0151] Among them, is the comprehensive priority of the th candidate command; is the speech clarity score; is the frequency of historical commands; is the relevance of the current 3D environment context; , and are weight coefficients, adjusted according to the importance of different factors.
[0152] Calculate the priorities of all candidate commands, and sort the candidate commands according to the priorities. The command with the highest priority will be selected as the next predicted command. Sorting formula:
[0153]
[0154] Among them, is the sorted set of candidate commands, arranged from high to low according to the priorities. represents the size of the candidate command set, that is, a total of candidate commands are generated; from the sorted set of candidate commands, select the command with the highest priority as the next predicted command.
[0155]
[0156] Among them, represents the prediction instruction for the next step, that is, the instruction with the highest priority. It means that after the candidate instructions are sorted by comprehensive scoring, the candidate instruction with the highest score is selected from them.
[0157] Preload relevant virtual resources for the prediction instruction of the next step, including 3D models, UI interfaces, sound effects, animations, etc. All these resources will be loaded according to the requirements of the current instruction.
[0158] For example, in a virtual power distribution scenario, the user vaguely issues an instruction "turn off that switch" through voice. The system identifies three candidate instructions, namely: turn off the main switch of distribution cabinet A1, turn off the standby switch of distribution cabinet B3, and turn off the console screen display.
[0159] For these three candidate instructions, the system comprehensively evaluates their corresponding speech clarity scores, historical instruction execution frequencies, and degrees of matching with the current environmental context. Among them, the speech clarity score for turning off the main switch of distribution cabinet A1 is 0.92, the historical execution frequency is 15 times, and the environmental context matching degree is 0.88; the clarity score for turning off the standby switch of distribution cabinet B3 is 0.78, the frequency is 3 times, and the matching degree is 0.65; the clarity score for turning off the console screen display is 0.85, the frequency is 1 time, and the matching degree is 0.40.
[0160] On the premise that the set weight coefficients are: speech clarity accounts for 0.4, historical frequency accounts for 0.3, and environmental matching degree accounts for 0.3, the system calculates the comprehensive priority of the three candidate instructions respectively. Finally, the comprehensive score of turning off the main switch of distribution cabinet A1 is the highest. The system takes it as the prediction instruction in the current context for priority response and preloads relevant execution resources in advance to achieve more efficient user instruction feedback.
[0161] By introducing a priority calculation method based on the relevance of speech clarity score, historical command execution frequency, and current 3D environmental context, the accuracy of candidate instruction generation and response has been significantly improved. The priority calculation method realizes the dynamic adjustment of priority by assigning different weight coefficients to the importance of each factor, enabling the system to better understand and predict user intentions. This priority calculation method based on multi-dimensional input makes the selection of candidate instructions more intelligent and can adapt to complex usage scenarios.
[0162] Furthermore, by preferentially loading virtual resources related to predicted instructions (such as 3D models, UI interfaces, sound effects, animations, etc.), the fluency and experience of user interaction are enhanced. Compared with traditional methods, dynamically optimizing the process of instruction generation and resource loading according to the real-time needs of users greatly improves the interaction efficiency and user experience in the virtual environment, especially in complex 3D environments and multi-task scenarios, with significant advantages.
[0163] Embodiment 2 is an embodiment of the present invention, which provides a voice interaction system based on 3D virtual reality, including:
[0164] A data module that receives the user's real-time voice, preprocesses it, establishes a speech recognition model, and analyzes the user's pronunciation features and clarity.
[0165] An optimization module that optimizes the speech recognition model based on the user's pronunciation features and clarity.
[0166] A prediction module that combines the user's historical commands and the current 3D environment context, and uses the optimized speech recognition model to generate a set of candidate predicted instructions for the user's next step; the set of candidate predicted instructions includes multiple candidate instructions.
[0167] A loading module that identifies the priority of each candidate instruction, takes the candidate instruction with the highest priority as the predicted instruction for the user's next step, preferentially loads the resources related to the predicted instruction, and quickly responds to the user's needs.
[0168] Embodiment 3 is an embodiment of the present invention, which is different from the previous two embodiments in that:
[0169] ]>If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0170] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0171] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0172] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0173] Example 4, an embodiment of the present invention, provides a voice interaction method and system based on 3D virtuality. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0174] The experiment simulated a typical 3D virtual operation scenario, and four users (numbered U01~U04) interacted in the virtual power equipment maintenance system respectively, issuing instructions such as switch, monitor, switchover, etc.
[0175] In the test preparation stage, first, the real-time voice input of users during simulated operations is collected. After Wiener filtering for noise reduction, its Mel Frequency Cepstral Coefficients (MFCCs) are extracted to form a time-frequency feature matrix with a frame length of 25 ms and an overlap of 10 ms. This matrix is input into a deep speech recognition model containing a CNN, max pooling, and fully connected structure. The high and low-frequency energies are extracted through wavelet transform to construct a clarity score. At the same time, the speech rate rhythm is analyzed in combination with a sliding window, and the pronunciation deviation degree is dynamically evaluated. A dual tolerance mechanism for speech rate and clarity is set up.
[0176] Subsequently, the historical instruction sequences of users are called, and the LSTM is used to extract the behavior path state, and the current context vector is constructed in combination with the real-time voice features. The system uses the fine-tuned meta-learning model to output the semantic-behavior joint embedding of each candidate instruction, and constructs a heterogeneous graph spectrum containing instruction nodes, object nodes, and context nodes. The graph structure is propagated through the GCN to obtain the graph structure semantic embedding of each candidate instruction, and the structural error is attributed and modeled in combination with the user's historical labels. Finally, the meta-model update path is adjusted with the structural attribution matrix $A$ to complete the adaptive fine-tuning of the user-specific rule engine.
[0177] The system finally outputs the personalized preference scores and graph structure semantic adaptation scores of each candidate instruction, which are fused into a comprehensive score. After sorting, the top M instructions are selected as the priority response set. The following are the main scoring results of 4 users in the experiment:
[0178] In the experiment, the speech clarity score of user U01 is 0.64, the speech rate feature score is 0.81, the historical instruction matching degree is 0.58, the graph structure adaptation score is 0.64, the personalized preference score is 0.72, and the final comprehensive score is 0.68.
[0179] In the experiment, the speech clarity score of user U02 is 0.93, the speech rate feature score is 0.78, the historical instruction matching degree is 0.54, the graph structure adaptation score is 0.78, the personalized preference score is 0.71, and the final comprehensive score is 0.74.
[0180] In the experiment, the speech clarity score of user U03 is 0.83, the speech rate feature score is 0.73, the historical instruction matching degree is 0.59, the graph structure adaptation score is 0.69, the personalized preference score is 0.75, and the final comprehensive score is 0.72.
[0181] In the experiment, the speech clarity score of user U04 is 0.86, the speech rate feature score is 0.46, the historical instruction matching degree is 0.74, the graph structure adaptation score is 0.70, the personalized preference score is 0.60, and the final comprehensive score is 0.65.
[0182] From the analysis of the experimental data, it can be seen that due to the high speech clarity score (0.93) and good structure adaptation score (0.78) of user U02, the final comprehensive score of the instruction reaches 0.74, and the system accurately infers the most appropriate operation intention in its current scenario. In contrast, although the speech clarity of user U04 reaches 0.86, due to its low speech rate score (0.46), there is a slight deviation in the model preference recognition, and the final comprehensive score is 0.65, lower than other users. This phenomenon indicates that in the low speech rate scenario, the system can still maintain the recognition accuracy for the speech rate adjustment mechanism, but the response score decreases slightly, verifying that the adaptive optimization mechanism of the system has stability but is still limited by the degree of input abnormality.
[0183] Compared with the existing recognition systems based only on speech or rule templates, the present invention forms a closed-loop logic chain from instruction generation to prediction feedback by introducing semantic-behavior joint feature modeling, GCN structure propagation and structure attribution adjustment path, not only improving the prediction accuracy, but also enhancing the fault tolerance for situations such as speech abnormality, fuzzy expression, and historical behavior jump. When facing inputs with fuzzy speech or abnormal speech rate, the recognition results of traditional models often deviate or lose key instruction information, while the multiple adjustment mechanisms and graph modeling framework embedded in the present invention can effectively locate structural uncertainties and complete real-time fine-tuning, effectively improving the accuracy of instruction prediction and the response speed of the interactive system.
[0184] In addition, the personalized rule engine of the present invention can complete rapid adaptation in the inner loop on the premise of only using 8-16 recent voice records of users, and fine-tune and update the path through structure attribution feedback in the outer loop, showing far stronger generalization ability than traditional speech recognition systems. The overall experimental results show that this method is particularly suitable for application environments that require rapid response and high-precision recognition in complex 3D virtual scenarios, and has strong adaptability, stability and promotion potential.
[0185] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A 3D virtual voice interaction method, characterized in that: include: Receive the user's real-time speech and pre-process it, build a speech recognition model, and analyze the user's pronunciation characteristics and clarity; Optimize the speech recognition model based on the user's pronunciation characteristics and clarity; Combining the user's historical commands and the current 3D environment context, using the optimized speech recognition model to generate a set of candidate predicted instructions for the user's next step; the candidate predicted instruction set includes multiple candidate instructions; Identify the priority of each candidate instruction, use the candidate instruction with the highest priority as the user's next predicted instruction, prioritize loading related resources of the predicted instruction, and quickly respond to user needs; identifying the priority of each candidate instruction includes calculating the priority of each candidate instruction based on the voice clarity score, the execution frequency of historical commands, and the matching degree of the instruction with the current environment.
2. The 3D virtual voice interaction method according to claim 1, wherein: The preprocessing includes using the Wiener filter algorithm to remove the noise of the real-time speech, dividing the real-time speech after denoising into multiple continuous frames of audio, and the overlapping part between each frame is ; Use Mel frequency cepstral coefficients to extract the time-frequency features of each frame of audio and establish a time-frequency feature matrix; where, Indicates the overlap length; The time-frequency features include the spectrum, pitch, volume, fundamental frequency and spectrum flatness of the audio.
3. The 3D virtual voice interaction method according to claim 2, wherein: The speech recognition model includes a convolutional neural network layer, a pooling layer and a fully connected layer; The time-frequency feature matrix is input into the convolutional neural network layer for convolution operation to extract local features and form a local feature map; the pooling layer uses the maximum pooling method to reduce the dimension of the local features to obtain the key local feature map; the fully connected layer converts the local feature map and the key local feature map into feature vectors to construct a global feature map.
4. The voice interaction method based on 3D virtuality according to claim 3, wherein: Analyzing the user's pronunciation clarity includes extracting high-frequency coefficients and low-frequency coefficients from the global feature map using a wavelet transform method; calculating the energy of the high-frequency coefficients and the energy of the low-frequency coefficients; determining a speech clarity score based on the ratio of the high-frequency energy to the low-frequency energy; Set the clarity score threshold. When the clarity score is lower than the clarity score threshold, it is determined that the speech clarity is insufficient, and the error tolerance rate is increased; Analyzing the user's pronunciation features includes using the sliding window method to calculate the local energy of high-frequency detail coefficients, finding the local energy peak, and determining the stressed part; calculating the average time interval between adjacent local energy peaks; setting the fast speech rate threshold and the slow speech rate threshold. When the average time interval is less than the fast speech rate threshold, it is determined that the speech rate is too fast, and the weight of the high-frequency filter in the convolutional neural network is increased to enhance the ability to capture rapidly changing syllables; when the average time interval is greater than the slow speech rate threshold, it is determined that the speech rate is too slow, and the weight of the low-frequency filter in the convolutional neural network is increased to enhance the recognition ability of slow-paced speech.
5. The voice interaction method based on 3D virtuality according to claim 4, wherein: Generating a set of candidate predicted instructions for the user's next step includes receiving a historical instruction sequence of the user; converting the historical instruction sequence into a time series feature using a long short-term memory network, generating a hidden state of the historical instructions, and recursively updating the hidden state to learn information in the historical instructions to obtain a historical command state; Utilize the optimized speech recognition model to analyze the received real-time speech input and obtain real-time speech features; Build a multi-layer deep model to deduce the user's possible operation intentions layer by layer and generate a set of candidate predicted instructions for the user's next step.
6. The voice interaction method based on 3D virtuality according to claim 5, wherein: The multi-layer deep model includes an LSTM layer, a GCN layer and an output layer; In the LSTM layer, the historical command status and real-time speech features are combined to generate a preliminary candidate command set using LSTM; Combining the historical command state with real-time speech features, and using LSTM to generate a real-time candidate command set; The GCN layer includes: based on the user's historical instruction sequence, pre-training the rule engine offline using a meta-learning framework to establish a user-specific rule model; when the user starts voice interaction, based on the current instruction speech clarity score and the average value of all syllable intervals in the current instruction, sampling a subset of historical interaction data with the same instruction speech clarity score and average value of all syllable intervals in the current instruction to construct a support set, and quickly fine-tuning the meta-model parameters to generate a current user-specific rule engine; for each candidate instruction, extracting a semantic-behavior joint feature vector using the support set; Based on the semantic-behavior joint feature vector, constructing a heterogeneous graph spectrum; The node set for constructing the heterogeneous graph spectrum includes: instruction nodes, target object nodes, and context semantic nodes; for any two nodes i, j in the heterogeneous graph spectrum, the corresponding semantic-behavior joint feature vectors are and When , an edge is established between i and j, and the weight of the edge is ; where represents the feature similarity, and represents the set structure connection threshold.
7. The voice interaction method based on 3D virtual as claimed in claim 6, wherein: Performing structure propagation on the heterogeneous graph spectrum through a multi-layer graph convolutional network to obtain a graph structure embedding representation of the candidate instruction, and outputting an instruction prediction result through a prediction head structure; comparing the prediction result with the true interaction behavior label to identify the prediction error, and when the prediction error exceeds a set threshold, extracting the structure risk factor of the candidate instruction and constructing a structure attribution matrix; The structure risk factor includes a semantic ambiguity factor and a path conflict factor; the semantic ambiguity factor is a quantitative index of the semantic concentration degree when the candidate instruction node points to multiple objects or semantic nodes: extracting all edges of the candidate instruction node and their corresponding edge weights, normalizing the edge weight set into a probability distribution, and then calculating the uncertainty degree of the probability distribution based on Shannon entropy, and taking the obtained entropy value after normalization as the value of the semantic ambiguity factor; The path conflict factor represents the degree of overlap between the path where the candidate instruction node is located and the adjacent candidate path in the graph structure on the structural edge set. The path is ,extract The edge set ; According to the figure and There are directly connected instruction nodes to construct neighbor path sets , extract each neighbor path separately The edge set , and calculate with The structural Jaccard similarity between the two is taken as the maximum value of the structural Jaccard similarity. ; Taking the structure attribution matrix as a regulatory factor and introducing it into the parameter update path of the personalized rule engine, and using the updated personalized rule engine to calculate the personalized preference score of each candidate instruction to guide the meta-model to avoid high-structure-risk areas in subsequent tasks; Fusing the personalized preference score and the semantic adaptation score to generate a comprehensive score for the candidate instruction, sorting the scores from high to low, and selecting the top M candidate instructions as the environmental context candidate instruction set; the semantic adaptation score is the embedding probability value output by the prediction head after the structure propagation of the graph convolutional network structure, which is used to measure the structural semantic fitting degree between the candidate instruction and the target object and context nodes in the graph; the output layer uses an attention mechanism to perform weighted fusion on the preliminary candidate instruction set, the real-time candidate instruction set, and the environmental context candidate instruction set to generate a candidate prediction instruction set for the user's next step.
8. A 3D virtual-based voice interaction system adopting the method according to any one of claims 1-7, characterized in that: A data module that receives the user's real-time voice, preprocesses it, establishes a speech recognition model, and analyzes the user's pronunciation features and clarity; An optimization module that optimizes the speech recognition model based on the user's pronunciation features and clarity; A prediction module that combines the user's historical commands and the current 3D environmental context, and uses the optimized speech recognition model to generate a candidate prediction instruction set for the user's next step; the candidate prediction instruction set includes multiple candidate instructions; The loading module identifies the priority of each candidate instruction, takes the candidate instruction with the highest priority as the predicted instruction for the user's next step, preferentially loads the relevant resources of the predicted instruction, and quickly responds to the user's needs.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the 3D virtual-based voice interaction method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the 3D virtual-based voice interaction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A voice interaction analysis method and system for an intelligent voice robot
CN119785777A
Voice instruction recognition method, system and device based on intention prediction and medium
CN119943034A
Device and method for voice activity detection
US20100088094A1
Voice Recognition Accuracy in High Noise Conditions
US20170256270A1
Cited By
Intelligent virtual accompanying robot scene dialogue rhythm control system
CN120656457A
Voice multi-instruction parallel recognition method and device based on large model
CN120877731A