Immersive virtual character interaction method based on virtual reality
By combining dynamic speech recognition and action recognition models, the user's voice and action information are processed and the performance status of virtual characters is adjusted, which solves the problem of insufficient leg and foot posture recognition and misunderstanding of speech recognition in virtual reality technology, and improves the user's immersion and interactive experience.
Patent Information
- Application Number
- CN202510156655.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In traditional virtual reality technology, the lack of recognition of postures such as legs and feet leads to the inability of users to effectively reflect their movements in virtual scenes, reducing the quality of the immersive experience. At the same time, the existing virtual reality human-computer interaction solutions have misunderstandings or errors in speech recognition, which is difficult to meet the individual differences and needs of different users.
The dynamic speech recognition model and action recognition model are used to obtain the user's voice information and action information, and the speech information is recognized through the dynamic speech recognition model. The action recognition model processes the action information, judges the correlation degree of speech and action information, adjusts the text and performance states to generate the target pronunciation and performance state of the virtual character, and optimizes the interaction information through semantic compensation methods.
It improves the accuracy and naturalness of the movement performance of virtual characters in virtual scenes, enhances the user's sense of immersion, solves the shortcomings of leg and foot posture recognition in traditional technology, and improves the accuracy of speech recognition to meet the individual different needs of different users.
Smart Images

Figure CN120085753A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and particularly to an immersive virtual character interaction method based on virtual reality. Background Art
[0002] Virtual Reality (VR) is a simulated environment generated by a computer, and users can immerse themselves in it and interact. With the development of virtual reality technology, it has been applied in many fields such as aerospace, medicine, architectural design, and education at home and abroad. Users experience and interact with objects in the virtual environment to obtain an immersive experience.
[0003] In traditional technologies, VR technology is also widely used in sports, medical treatment, and education. In sports and medical treatment, users are usually trained in a preset manner. However, in the field of virtual scenarios, since most VR devices (such as HTC Vive) can only be operated with a handle and lack the recognition of postures such as legs and feet, the actions made by users cannot be reflected in the virtual scenario, resulting in a poor immersive experience for users. In addition, when users interact with virtual characters through voice, in existing virtual reality human-computer interaction solutions, the speech recognition technology may have misunderstandings or errors when recognizing users' voice commands, which may cause digital characters to fail to correctly understand users' intentions, and thus unable to make correct actions and dialogue effects. In addition, existing virtual reality human-computer interaction solutions may not be able to meet the individual differences and needs of different users. For example, for users with accents or different language habits, the speech recognition technology may have difficulties, resulting in poor interaction effects.
[0004] Therefore, the present invention provides an immersive virtual character interaction method based on virtual reality to solve the above problems. Summary of the Invention
[0005] In view of the above situation, to overcome the deficiencies of the prior art, the present invention provides an immersive virtual character interaction method based on virtual reality to solve the problem that the above traditional virtual interaction method lacks the recognition of postures such as legs and feet, so that the actions made by users cannot be reflected in the virtual scenario, resulting in a poor immersive experience for users.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] An immersive virtual character interaction method based on virtual reality, comprising: obtaining voice information input by a user and action information collected by an action capture device;
[0008] Use a dynamic speech recognition model to recognize speech information to obtain initial text information and speech feature information; and use an action recognition model to process action information to determine initial performance state information composed of the action features of the virtual character;
[0009] Judge whether the degree of association between the initial text information and the initial performance state information meets the preset association degree standard; if so, adjust the initial text information of the user according to the initial performance state information of the user to obtain target text information, and generate target pronunciation information of the virtual character according to the target text information and the speech feature information; and adjust the initial performance state information according to the target text information and the speech feature information to obtain the target performance state information of the virtual character; if not, use a semantic compensation method to compensate the initial text information to obtain target text information, generate target pronunciation information of the virtual character according to the target text information and the speech feature information; and use the initial performance state information as the target performance state information;
[0010] Based on the current interaction scenario, fuse the target pronunciation information and the target performance state information to form the interaction information of the virtual character.
[0011] Preferably, the using a dynamic speech recognition model to recognize speech information to obtain speech feature information and initial text information includes: using an encoder to preprocess the speech information, converting the preprocessed speech information into a plurality of spectral feature maps, processing each spectral feature map to obtain speech feature information and a sequence of encoded vectors; processing the sequence of encoded vectors based on an attention mechanism to obtain attention weights at each position, and using a weighted average method to process the sequence of encoded vectors and the attention weights to obtain a key vector; using a decoder to analyze the sequence of encoded vectors and the key vector to obtain initial text information.
[0012] Preferably, the using an action recognition model to process action information to determine initial performance state information composed of the action features of the virtual character includes: adjusting the actions of the initial virtual character according to the action information to obtain a first performance action; adjusting the first performance action according to the setting criteria of the initial virtual character corresponding to each attribute to obtain a second performance action that meets the setting criteria of the initial virtual character; adding corresponding appearance information to the second performance action to obtain the initial performance state information of the virtual character.
[0013] Preferably, the judging whether the degree of association between the initial text information and the initial performance state information meets the preset association degree standard includes: determining a semantic intentionality index score according to the initial text information and the initial performance state information, and the semantic intentionality index score S in The calculation formula of is:
[0014]
[0015] Among them, a i is the i-th word vector in the word feature vector set corresponding to the initial text information, and b i is the i-th word vector in the word feature vector set corresponding to the initial performance state information. The total number of word vectors in the word feature vector set is n;
[0016] Determine the scene relevance index score S according to the scene requirement rules, the initial text information, and the initial performance state information sc , S sc The calculation formula is:
[0017] S sc = ω 1 * T + ω 2 * P + ω 3 * R,
[0018] Among them, ω 1 , ω 2 , ω 3 are the weight coefficients corresponding to T, P, and R respectively. T is the text relevance score, P is the performance suitability score, and R is the requirement compliance score;
[0019] Determine the historical behavior relevance index score S according to the historical behavior information, the initial text information, and the initial performance state information hi , S hi The calculation formula of is:
[0020] S hi = ω 4 * Pa(x c0i , x c1i ) + ω 5 * Te(y e0i , y e1i ) + ω 6 * Pe,
[0021] Among them, ω 4 , ω 5 , ω 6 are the weight coefficients corresponding to Pa(x c0i , x c1i ), Te(y e0i , y e1i ), and Pe respectively. Pa(x c0i , x c1i ) is the past behavior similarity scoring function, x c0i is the i-th feature vector in the historical behavior information feature vector set, 0 is the historical marker symbol, c is the scene type or behavior pattern type, x c1iis the i-th feature vector in the behavioral information feature vector set of the initial text information and the initial performance state information. 1 is the current token, Te(y e0i ,y e1i ), is the text-based relevance scoring function, y e0i is the i-th word vector in the vector set formed by the text content in the historical behavioral information, e0 is the historical text token, y e1i is the i-th word vector in the vector set formed by the initial text information, e1 is the token of the current text information, Pe is the performance consistency score;
[0022] When the qualified quantity of each index score is greater than the preset qualified quantity, it is determined that the preset correlation degree standard is met.
[0023] Preferably, adjusting the initial text information of the user according to the initial performance state information of the user to obtain the target text information includes: parsing the initial performance state information to obtain key pose features; analyzing the key pose features based on a preset pose type discrimination model to determine the pose category; determining the corresponding text adjustment strategy from the mapping relationship set according to the pose category; adjusting the initial text information according to the text adjustment strategy to obtain the target text information.
[0024] Preferably, adjusting the initial performance state information according to the target text information and the voice feature information to obtain the target performance state information of the virtual character includes: performing semantic analysis on the target text information to determine the first theme feature of the target text; performing feature analysis on the voice feature information to determine the second theme feature of the voice feature; fusing the first theme feature and the second theme feature to obtain the fused theme feature; searching for the corresponding action adjustment strategy from the fused adjustment rule library according to the fused theme feature; adjusting the initial performance state information according to the action adjustment strategy to obtain the target performance state information.
[0025] Preferably, compensating the initial text information by using a semantic compensation method to obtain the target text information includes: performing semantic analysis on the initial text information to obtain the semantic analysis result; determining the missing position and the semantic missing keyword in the initial text information according to the semantic analysis result and the initial text information; querying the compensation information corresponding to the semantic missing keyword from the semantic compensation knowledge base; performing semantic compensation on the initial text information at the missing position according to the compensation information to obtain the target text information.
[0026] Preferably, fusing the target pronunciation information and the target performance state information based on the current interaction scenario to form the interaction information of the virtual character includes: determining the pronunciation type of the virtual character according to the current interaction scenario; adjusting the target pronunciation information according to the pronunciation type; synchronously integrating the adjusted target pronunciation information and the target performance state information in the order of pronunciation and action to form the real-time interaction information of the virtual character.
[0027] Preferably, an immersive virtual character interaction method based on virtual reality further includes: when the user creates a virtual character, obtaining the sample voice information input by the user, using a preset voice recognition model to recognize the sample voice information, extracting the sample voice features, and adjusting the preset sample voice feature library according to the sample voice features to obtain the initial voice feature library of the user.
[0028] The dynamic voice recognition model further includes: after obtaining the voice feature information, based on the voice feature information, determining whether the initial voice feature library contains unadded voice features, and if so, adding the unadded voice features to the corresponding positions in the dynamic voice feature library to obtain the dynamic voice feature library.
[0029] The beneficial effects of the present invention are as follows:
[0030] 1. The present invention can adjust the initial performance state information on the basis of the initial performance state information and when the correlation degree between the initial text information and the initial performance state information meets the preset correlation degree standard, so that the actions shown by the virtual character are more in line with the user's needs, thereby increasing the user's immersion. And the present invention solves the problem that the traditional virtual interaction method lacks the recognition of postures such as legs and feet, so that the actions made by the user cannot be reflected in the virtual scene, resulting in poor immersive experience of the user, through the above method.
[0031] 2. The present invention adopts the method of processing the action information collected by the motion capture device to obtain the initial performance state information composed of the action features of the virtual character; and on the basis of the voice feature information and the initial text information obtained by analyzing the voice information input by the user, determining whether the correlation degree between the initial text information and the initial performance state information meets the preset correlation degree standard; when the correlation degree meets the preset correlation degree standard, adjusting the initial performance state information according to the target text information and the voice feature information to obtain the target performance state information of the virtual character, and adjusting the initial text information of the user according to the initial performance state information of the user to obtain the target text information, and then being able to combine the target performance state information and the target text information into the interaction information shown by the virtual character. Through the above method, the present invention can make the interaction voice and actions shown by the virtual character more in line with the user's needs, thereby increasing the user's immersion.
[0032] 3. When determining whether the correlation degree between the initial text information and the initial presentation state information meets the preset correlation degree standard, the present invention judges the correlation degree from three indicators: semantic intention, scenario relevance, and historical behavior relevance. When the number of qualified scores of each indicator is greater than the preset qualified number, it is determined that the preset correlation degree standard is met. Through the above judgment method, the present invention can accurately judge whether the initial text information and the initial presentation state information are related; and then execute different optimization methods to process the initial text information and the initial presentation state information, so that the pronunciation information and actions shown by the user's virtual character conform to the user's needs and the actual situation of the current scenario, thereby helping to increase the user's immersion. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of an immersive virtual character interaction method based on virtual reality according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, each embodiment of the present invention will be described in detail with reference to the attached Figure 1 Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.
[0035] An immersive virtual character interaction method based on virtual reality, as shown in the attached Figure 1 figure, includes the following steps:
[0036] Step S11: Obtain the voice information input by the user and the action information collected by the motion capture device, where the action information includes hand actions, foot actions, and eye actions.
[0037] Among them, hand actions are recognized through gloves or handles, such as scissors, rock, paper, etc.; foot actions are captured by foot capture devices, such as the tilt degree and lifting height of the feet; eye actions such as eyeball rotation, eyelid movement, and eyebrow center state of the user are captured through VR glasses or head-mounted devices.
[0038] Step S12: Use a dynamic speech recognition model to recognize the voice information to obtain voice feature information and initial text information.
[0039] Among them, the voice feature information includes the sound intensity, decibel, speech rate, zero-crossing rate, fundamental frequency and pitch period, formant, short-time average energy and short-time average amplitude, emotion information, and language information of the voice.
[0040] Step S13: Use an action recognition model to process the action information to determine the initial presentation state information composed of the action features of the virtual character.
[0041] Among them, the initial performance state information is obtained by fusing various action features.
[0042] Step S14: Determine whether the correlation degree between the initial text information and the initial performance state information meets the preset correlation degree standard. If so, execute Step S15; if not, execute Step S16.
[0043] Among them, the preset correlation degree standard includes semantic intentionality, scene relevance, and historical behavior relevance.
[0044] Step S15: When the correlation degree meets the preset correlation degree standard, adjust the initial text information of the user according to the initial performance state information of the user to obtain the target text information, and generate the target pronunciation information of the virtual character according to the target text information and the voice feature information; and adjust the initial performance state information according to the target text information and the voice feature information to obtain the target performance state information of the virtual character.
[0045] Step S16: When the correlation degree does not meet the preset correlation degree standard, compensate the initial text information by using a semantic compensation method to obtain the target text information, generate the target pronunciation information of the virtual character according to the target text information and the voice feature information; and use the initial performance state information as the target performance state information.
[0046] Step S17: Based on the current interaction scenario, fuse the target pronunciation information and the target performance state information to form the interaction information of the virtual character.
[0047] In this embodiment, the present invention can adjust the initial performance state information on the basis of the initial performance state information and when the correlation degree between the initial text information and the initial performance state information meets the preset correlation degree standard, so that the actions shown by the virtual character are more in line with the user's needs, thereby increasing the user's immersion. Furthermore, the present invention solves the problem that the traditional virtual interaction method lacks the recognition of postures such as legs and feet, so that the actions made by the user cannot be reflected in the virtual scene, resulting in a poor immersive experience for the user, through the above method.
[0048] Specifically, the present invention processes the motion information collected by a motion capture device to obtain initial presentation state information constituted by the motion characteristics of a virtual character; and based on the speech feature information and initial text information obtained by analyzing the speech information input by a user, determines whether the degree of association between the initial text information and the initial presentation state information meets a preset degree of association standard; when the degree of association meets the preset degree of association standard, adjusts the initial presentation state information according to the target text information and the speech feature information to obtain the target presentation state information of the virtual character, and adjusts the initial text information of the user according to the initial presentation state information of the user to obtain the target text information, and further can combine the target presentation state information and the target text information into interaction information presented by the virtual character. Through the above method, the present invention can make the interaction speech and actions presented by the virtual character natural and smooth in line with the user's needs, thereby increasing the user's immersion.
[0049] In one embodiment of the present invention, the step of using a dynamic speech recognition model to recognize speech information to obtain speech feature information and initial text information includes: preprocessing the speech information by using an encoder, converting the preprocessed speech information into a plurality of spectral feature maps, processing each spectral feature map to obtain speech feature information and a sequence of encoded vector; processing the sequence of encoded vector based on an attention mechanism to obtain attention weights at each position, and using a weighted average method to process the sequence of encoded vector and the attention weights to obtain a key vector; analyzing the sequence of encoded vector and the key vector by using a decoder to obtain the initial text information.
[0050] In this embodiment, the encoder stage includes: Input processing and feature extraction: First, preprocess the input speech signal, such as sampling, filtering, etc., and then convert the speech signal into a spectral map or a Mel spectral map and other feature representations through methods such as short-time Fourier transform. These feature maps are used as the input of the encoder. The encoder generally adopts a structure such as a convolutional neural network, and performs convolutional and pooling operations on the feature maps to extract the local features and global information of the speech signal. Generating a hidden representation: The encoder encodes the input sequence and maps it into a sequence of hidden representation vectors of a fixed length, that is, a sequence of encoded vectors. This hidden representation contains important information of the input speech signal, but removes the noise and redundant information in the original signal, making subsequent processing more efficient and accurate.
[0051] The attention mechanism stage includes: Calculating attention weights: For each position in the input sequence of encoded vectors, calculate the attention weights for each position through a fully connected layer and a Softmax activation function. Specifically, first perform a dot product operation on the sequence of encoded vectors output by the encoder and a learnable query vector to obtain the scores for each position, and then normalize them through the Softmax function to obtain the attention weights for each position. Weighted summation to generate the key vector: Multiply each position in the input sequence by the corresponding attention weight, and then sum up the weighted results of all positions to obtain a key vector of a fixed length. This key vector can be regarded as a weighted summary of the key information in the input sequence of encoded vectors, which can highlight the information most helpful for the current output.
[0052] The decoder stage includes: Initialization and iterative prediction: The decoder usually adopts structures such as recurrent neural networks (RNNs) or long short-term memory networks (LSTMs), takes the sequence of encoded vectors output by the encoder as the initial input, and combines the key vector generated by the attention mechanism to start predicting the output sequence frame by frame or word by word. In each step of prediction, the decoder calculates the probability distribution of the next output based on the current encoded vector and the key vector, and then selects the output with the highest probability as the prediction result for the current frame or the current word. Updating the key vector and the encoded vector: The decoder updates the sequence of encoded vectors and the key vector according to the current output prediction result, so as to better utilize the previous information and the newly generated output information in the next step of prediction. The output prediction results are combined in order to form the initial text information. At the same time, the attention mechanism also recalculates the attention weights according to the new sequence of encoded vectors, enabling the model to dynamically focus on different parts of the sequence of encoded vectors.
[0053] The training and optimization stage includes: Loss calculation: Compare the predicted output sequence with the true target sequence to calculate the loss value. Commonly used loss functions include the cross-entropy loss function, etc. The loss function is used to measure the difference between the model prediction result and the true result. Backpropagation and parameter update: Use the backpropagation algorithm to pass the loss value to each parameter of the model, and adjust and update the parameters of the model according to the loss value. Through continuous iterative training, the model gradually learns the mapping relationship from the input speech signal to the output text sequence, improving the accuracy of speech recognition.
[0054] In one embodiment of the present invention, processing the action information using an action recognition model to determine the initial presentation state information constituted by the action characteristics of the virtual character includes: adjusting the actions of the initial virtual character according to the action information to obtain the first presentation action; adjusting the first presentation action according to the setting criteria of the initial virtual character corresponding to each attribute to obtain the second presentation action that conforms to the setting criteria of the initial virtual character; and adding corresponding appearance information to the second presentation action to obtain the initial presentation state information of the virtual character.
[0055] Preferably, the action rationality criteria include the deviation angle between the head and the body, the deviation angle between the feet and the calves, the length of the legs, the positions and bending angles of each finger, the change range of facial expression actions, etc.
[0056] Preferably, the appearance information includes hairstyle, clothing, shoes, accessories, tools, weapons, or other items, etc.
[0057] In this embodiment, the initial virtual character is a virtual character template without any modifiers, and the first presentation action determined according to the acquired action information can be intuitively seen; then, the action rationality criteria of the initial virtual character are obtained according to the attributes of the initial virtual character selected by the user; and the first presentation action of the initial virtual character is adjusted according to the action rationality criteria of the initial virtual character to obtain the second presentation action that conforms to the action rationality criteria; then, adaptively adjusted appearance information is added to the initial virtual character under the second virtual action. The attributes include the ability to stretch the body at will, the body can turn into water, the body can petrify, etc. Through the setting method of this embodiment, the present invention can ensure the rationality of the actions of the virtual character and make the actions of the virtual character conform to the action rationality criteria under the preset attributes.
[0058] In one embodiment of the present invention, determining whether the degree of association between the initial text information and the initial presentation state information meets the preset degree of association criteria includes: determining the semantic intention index score according to the initial text information and the initial presentation state information, and the semantic intention index score S in The calculation formula of is:
[0059]
[0060] Among them, a i Is the i-th word vector in the word feature vector set corresponding to the initial text information, and b i Is the i-th word vector in the word feature vector set corresponding to the initial presentation state information, and the total number of word vectors in the word feature vector set is n.
[0061] Determine the scene relevance index score S according to the scene requirement rules, the initial text information and the initial presentation state information sc ,Ssc The calculation formula is:
[0062] S sc = ω 1 *T + ω 2 *P + ω 3 *R,
[0063] where ω 1 , ω 2 , ω 3 are the weight coefficients corresponding to T, P, and R respectively. T is the text relevance score, P is the performance suitability score, and R is the requirement compliance score. The text relevance score is determined based on the initial text information and the scenario theme. The performance suitability score is determined based on the initial performance state information and the scenario theme, and is determined according to the rules of whether the initial text information and the initial performance state information meet the scenario requirements.
[0064] Determine the historical behavior relevance index score S according to the historical behavior information, the initial text information, and the initial performance state information hi S hi The calculation formula is:
[0065] S hi = ω 4 *Pa(x c0i , x c1i ) + ω 5 *Te(y e0i , y e1i ) + ω 6 *Pe,
[0066] where ω 4 , ω 5 , ω 6 are the weight coefficients corresponding to Pa(x c0i , x c1i ), Te(y e0i , y e1i ), and Pe respectively. Pa(x c0i , x c1i ) is the past behavior similarity scoring function, x c0i is the i-th feature vector in the historical behavior information feature vector set, 0 is the historical marker symbol, c is the scenario type or behavior pattern type, x c1i is the i-th feature vector in the behavior information feature vector set of both the initial text information and the initial performance state information, 1 is the current marker symbol, Te(y e0i , y e1i ) is the text-based relevance scoring function, y e0i is the i-th word vector in the vector set formed by the text content in the historical behavior information, e0 is the historical text marker symbol, ye1i It is the i-th word vector in the vector set formed by the initial text information, e1 is the token symbol of the current text information, and Pe is the performance consistency score.
[0067] Among them, the past behavior similarity score is determined based on the similarity degree between historical behavior information in different scenarios and different behavior patterns and the current initial text information and the current initial performance state information; the text-based relevance score is determined based on the relevant text content in the initial text information and historical behavior information; the performance consistency score is determined by whether the initial performance state information is consistent or coherent with the performance state information in historical behavior.
[0068] When the qualified quantity of each index score is greater than the preset qualified quantity, it is determined that the preset correlation degree standard is met.
[0069] Preferably, the preset correlation degree standard includes the evaluation rules of each preset correlation index. After quantifying the corresponding information according to the evaluation rules, the index score is obtained. The passing line of the index score is 60, and the total score is 100. The preset qualified quantity is that at least two index scores are qualified.
[0070] The preset correlation indexes include semantic intention index, scene relevance index, and historical behavior relevance index.
[0071] Through the setting method of this embodiment, the present invention can calculate the index scores of each index through the above calculation formula, so as to judge whether the correlation degree between the initial text information and the initial performance state information meets the preset correlation degree standard according to each index score, and then execute different subsequent steps to obtain the target performance state information of the virtual character, making the performance of the virtual character fit the user's needs, increasing the interaction accuracy and naturalness of the user's interaction with the virtual character, and further helping to improve the user's immersion, allowing the user to truly enjoy the fun brought by the virtual character. The present invention can accurately judge whether the initial text information and the initial performance state information are related based on the scoring conditions of the above indexes; and then execute different optimization methods to process the initial text information and the initial performance state information, so that the pronunciation information and actions shown by the user's virtual character fit the user's needs and conform to the actual situation of the current scene, thereby increasing the user's immersion.
[0072] In an embodiment of the present invention, adjusting the user's initial text information according to the user's initial performance state information to obtain the target text information includes: parsing the initial performance state information to obtain key pose features; analyzing the key pose features based on a preset pose type discrimination model to determine the pose category; determining the corresponding text adjustment strategy from the mapping relationship set according to the pose category; and adjusting the initial text information according to the text adjustment strategy to obtain the target text information.
[0073] Through the setting method of this embodiment, the present invention can adjust the initial text information of the user according to the initial performance state information of the user to obtain the target text information, so that the pronunciation information transformed from the target text information is more in line with the pronunciation tone and pronunciation content of the virtual character in the current scene, thereby helping to increase the user's fun and increase the user's immersion time.
[0074] In an embodiment of the present invention, adjusting the initial performance state information according to the target text information and the speech feature information to obtain the target performance state information of the virtual character includes: performing semantic analysis on the target text information to determine the first theme feature of the target text; performing feature analysis on the speech feature information to determine the second theme feature of the speech feature; fusing the first theme feature and the second theme feature to obtain a fused theme feature; searching for a corresponding action adjustment strategy from the fused adjustment rule library according to the fused theme feature; and adjusting the initial performance state information according to the action adjustment strategy to obtain the target performance state information.
[0075] Among them, the first theme feature includes the emotional tendency (positive, negative or neutral), theme content, etc. of the target text information; the second theme feature includes the emotional state expressed by the speech, for example, excited, depressed, upset, etc. The fused adjustment rule library records the mapping relationship between the fused theme feature and the action adjustment strategy, and can find the corresponding action adjustment strategy from the fused adjustment rule library according to the fused theme feature.
[0076] Through the setting method of this embodiment, the present invention can adjust the initial performance state information according to the target text information and the speech feature information to obtain the target performance state information of the virtual character, so that the performance actions of the virtual character meet the requirements in the current scene, making the performance actions of the virtual character more natural and smooth, thereby helping to increase the user's immersion.
[0077] In an embodiment of the present invention, compensating the initial text information by using a semantic compensation method to obtain the target text information includes: performing semantic analysis on the initial text information to obtain a semantic analysis result; determining the missing position and semantic missing keywords in the initial text information according to the semantic analysis result and the initial text information; querying the compensation information corresponding to the semantic missing keywords from the semantic compensation knowledge base; and performing semantic compensation on the initial text information at the missing position according to the compensation information to obtain the target text information.
[0078] In the above manner, when the correlation degree between the initial text information and the initial presentation state information does not meet the preset correlation degree standard, the present invention can use a semantic compensation method to compensate the initial text information to obtain target text information, thereby making the pronunciation information of the virtual character more accurate and more in line with the actual situation in the scene where the virtual character is located, enabling the user of the other virtual character in the interaction to more clearly understand the intention of the present user; and further helping to improve the user's immersion sense.
[0079] In an embodiment of the present invention, the method of fusing the target pronunciation information and the target presentation state information based on the current interaction scene to form the interaction information of the virtual character includes: determining the pronunciation type of the virtual character according to the current interaction scene; adjusting the target pronunciation information according to the pronunciation type; synchronously integrating the adjusted target pronunciation information and the target presentation state information in the order of pronunciation and action to form the real-time interaction information of the virtual character.
[0080] In this embodiment, the pronunciation mode of the user's virtual character is determined according to the decibel of the user's voice, the user's operation type, and the current scene characteristics; in the dynamic voice feature library, the corresponding voice features are specifically queried according to the pronunciation mode. When updating the dynamic voice feature library, the voice feature information in different interaction modes is collected and updated, so that the corresponding voice features can be specifically queried during the process of determining the interaction information of the virtual character, and the target pronunciation information can be quickly determined, which helps to improve the data processing speed and further helps to improve the user's experience.
[0081] In an embodiment of the present invention, a virtual reality-based immersive virtual character interaction method further includes: when the user creates a virtual character, obtaining the sample voice information input by the user, using a preset voice recognition model to recognize the sample voice information, extracting the sample voice features, and adjusting the preset sample voice feature library according to the sample voice features to obtain the user's initial voice feature library. And, in the case where the user does not input sample voice information, using the preset sample voice feature library as the user's initial voice feature library.
[0082] Through the setting method of this embodiment, the present invention can avoid the cold start problem, allowing the user to experience the fun of different scenes when just creating a virtual character; as the user uses the virtual character for a longer time, the initial voice feature library will gradually form the user's personal dynamic voice feature library, fully reflecting the user's actual control behavior on the virtual character, which helps to increase the user's immersion time and allows the user to experience different life pleasures on the virtual character.
[0083] In an embodiment of the present invention, the dynamic speech recognition model further includes: after obtaining the speech feature information, based on the dynamic speech feature library, determining whether the current speech feature information contains unadded speech features. If so, adding the unadded speech features to the corresponding positions in the dynamic speech feature library to update the dynamic speech feature library.
[0084] Through the setting method of this embodiment, the present invention can generate a dynamic speech feature library unique to the user during the process of the user experiencing the life of the virtual character for a long time. Furthermore, when obtaining the user's speech information again, the target pronunciation information of the virtual character can be quickly generated through the dynamic speech recognition model, which helps to improve the immersion and dependence of the user experiencing the virtual character.
[0085] Furthermore, when establishing a virtual character, there are three methods. The first method: select a target character template from the preset virtual character templates, obtain the adjustment information of the target character template by the user using the preset character adjustment tool and display it in real time to obtain the user's virtual character; the second method: the user customizes the virtual character by painting; the third method: generate the initial virtual character of the user based on the picture information uploaded by the user, obtain the adjustment information of the initial virtual character by the user using the preset character adjustment tool and display it in real time to obtain the user's virtual character.
[0086] In an embodiment of the present invention, during the process of updating the feature information, the replaced speech feature information and action feature information are compressed and backed up so as to generate the memory experience of the user during the use of the virtual character during the anniversary event, increasing the user's dependence and experience.
[0087] In an embodiment of the present invention, an advanced preset lighting model and texture mapping model are used to render the virtual character, increasing the three-dimensional sense and external appearance image of the virtual character, such as facial expression actions, etc.
[0088] In one embodiment of the present invention, to improve the interaction accuracy and naturalness: enhance the research and development of input devices and interaction technologies, and improve the accuracy and speed of technologies such as gesture recognition, speech recognition, and eye tracking. Adopt more advanced sensors and algorithms to capture the user's actions and intentions more accurately, and achieve more natural and smooth human-computer interaction. To reduce network latency and data transmission errors, optimize the network architecture, ensure real-time synchronization and interaction consistency among multiple users, establish effective collaborative working mechanisms and interaction rules, enable different users to better cooperate to complete tasks, and enhance the experience and effect of teamwork. And optimize hardware devices and technologies: continuously research and develop more advanced hardware devices, improve processing capabilities and graphics performance, reduce device costs and volumes, so as to enhance the user's usage experience and comfort. At the same time, improve software algorithms and optimize technologies to reduce latency and stuttering phenomena, and improve the fluency and stability of interaction.
[0089] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0091] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0092] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0093] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0094] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0095] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. An immersive virtual character interaction method based on virtual reality, characterized in that: include: Acquire voice information input by the user and motion information collected by the motion capture device; A dynamic speech recognition model is used to recognize speech information to obtain initial text information and speech feature information; and processing the action information using an action recognition model to determine initial performance state information consisting of action features of the virtual character; Determine whether the correlation between the initial text information and the initial performance state information meets the preset correlation degree standard; if so, adjust the initial text information of the user according to the initial performance state information of the user to obtain the target text information, and generate the target pronunciation information of the virtual character according to the target text information and the voice feature information; and adjust the initial performance state information according to the target text information and the voice feature information to obtain the target performance state information of the virtual character; if not, compensate the initial text information by using a semantic compensation method to obtain the target text information, and generate the target pronunciation information of the virtual character according to the target text information and the voice feature information; and use the initial performance state information as the target performance state information; Based on the current interaction scenario, the target pronunciation information and the target performance status information are fused to form the interaction information of the virtual character.
2. The immersive virtual character interaction method according to claim 1, characterized in that: The method of using a dynamic speech recognition model to recognize speech information to obtain speech feature information and initial text information includes: using an encoder to preprocess the speech information, converting the preprocessed speech information into multiple spectrum feature maps, and processing each spectrum feature map to obtain speech feature information and a coding vector sequence; processing the coding vector sequence based on an attention mechanism to obtain an attention weight for each position, and using a weighted average method to process the coding vector sequence and the attention weight to obtain a key vector; and using a decoder to analyze the coding vector sequence and the key vector to obtain initial text information.
3. The immersive virtual character interaction method according to claim 1, characterized in that: The method of using an action recognition model to process action information to determine initial performance status information composed of action features of a virtual character includes: adjusting the action of the initial virtual character according to the action information to obtain a first performance action; adjusting the first performance action according to the setting standards of the initial virtual character corresponding to each attribute to obtain a second performance action that meets the setting standards of the initial virtual character; and adding corresponding appearance information to the second performance action to obtain the initial performance status information of the virtual character.
4. The immersive virtual character interaction method according to claim 3, characterized in that: The determination of whether the correlation between the initial text information and the initial presentation status information meets the preset correlation degree standard includes: determining a semantic intention index score based on the initial text information and the initial presentation status information, wherein the semantic intention index score S in The calculation formula is: Among them, a i is the i-th word vector in the word feature vector set corresponding to the initial text information, b i is the i-th word vector in the word feature vector set corresponding to the initial performance state information, and the total number of word vectors in the word feature vector set is n; Determine the scene relevance index score S based on the scene requirement rules, initial text information and initial performance status information sc , S sc The calculation formula is: S sc =ω1*T+ω2*P+ω3*R, Among them, ω1, ω2, and ω3 are the weight coefficients corresponding to T, P, and R respectively, T is the text relevance score, P is the performance suitability score, and R is the requirement compliance score; Determine the historical behavior relevance index score S based on the historical behavior information, initial text information, and initial performance status information hi , S hi The calculation formula is: S hi =ω4*Pa(x c0i ,x c1i )+ω5*Te(y e0i ,y e1i )+ω6*Pe, Among them, ω4, ω5, and ω6 are Pa(x c0i ,x c1i )、Te(y e0i ,y e1i ), the weight coefficient corresponding to Pe, Pa(x c0i ,x c1i ) is the past behavior similarity scoring function, x c0i is the i-th feature vector in the historical behavior information feature vector set, 0 is the historical mark symbol, c is the scene type or behavior pattern type, x c1i is the i-th feature vector in the feature vector set of the behavior information in the initial text information and the initial performance state information, 1 is the current marking symbol, Te(y e0i ,y e1i ) is the text-based relevance scoring function, y e0i is the i-th word vector in the vector set formed by the text content in the historical behavior information, e0 is the historical text mark symbol, y e1i is the i-th word vector in the vector set formed by the initial text information, e1 is the marker symbol of the current text information, and Pe is the performance consistency score; When the number of qualified scores for each indicator is greater than the preset qualified number, it is judged to meet the preset correlation degree standard.
5. The immersive virtual character interaction method according to claim 1, characterized in that: The method of adjusting the user's initial text information according to the user's initial performance status information to obtain target text information includes: parsing the initial performance status information to obtain key posture features; analyzing the key posture features based on a preset posture type discrimination model to determine the posture category; determining the corresponding text adjustment strategy from a mapping relationship set according to the posture category; and adjusting the initial text information according to the text adjustment strategy to obtain the target text information.
6. The immersive virtual character interaction method according to claim 1, characterized in that: The method of adjusting the initial performance state information according to the target text information and the voice feature information to obtain the target performance state information of the virtual character includes: performing semantic analysis on the target text information to determine the first theme feature of the target text; performing feature analysis on the voice feature information to determine the second theme feature of the voice feature; fusing the first theme feature and the second theme feature to obtain a fused theme feature; searching for a corresponding action adjustment strategy from a fused adjustment rule library according to the fused theme feature; and adjusting the initial performance state information according to the action adjustment strategy to obtain the target performance state information.
7. The immersive virtual character interaction method according to claim 1, characterized in that: The method of using semantic compensation to compensate initial text information to obtain target text information includes: performing semantic analysis on the initial text information to obtain semantic analysis results; determining missing positions and semantic missing keywords in the initial text information based on the semantic analysis results and the initial text information; querying compensation information corresponding to the semantic missing keywords from a semantic compensation knowledge base; and performing semantic compensation on the initial text information at the missing positions based on the compensation information to obtain the target text information.
8. The immersive virtual character interaction method according to claim 7, characterized in that: The method of fusing the target pronunciation information and the target performance status information based on the current interaction scenario to form the interaction information of the virtual character includes: According to the current interaction scenario, the pronunciation type of the virtual character is determined; the target pronunciation information is adjusted according to the pronunciation type; the adjusted target pronunciation information and target performance status information are synchronously integrated according to the sequence of pronunciation and action to form real-time interaction information of the virtual character.
9. The immersive virtual character interaction method according to claim 2, characterized in that: Also includes: When a user creates a virtual character, sample voice information input by the user is obtained, the sample voice information is recognized using a preset voice recognition model, sample voice features are extracted, and a preset sample voice feature library is adjusted according to the sample voice features to obtain the user's initial voice feature library.
10. The immersive virtual character interaction method according to claim 9, characterized in that: The dynamic speech recognition model also includes: after obtaining the speech feature information, based on the speech feature information, determining whether the initial speech feature library contains unadded speech features, and if so, adding the unadded speech features to the corresponding positions in the dynamic speech feature library to obtain the dynamic speech feature library.
Citation Information
Cited By
Digital human image generation method combining cold start driving and active learning mechanism
CN120833401A
Video AR virtual reality interactive communication method and system based on AI model
CN121788771A
AI model-based video AR virtual reality interactive communication method and system
CN121788771B