Virtual character model with voice control function and control method thereof

By timbre identification and feature extraction of voice control commands in multi-person voice environments, combined with matching and analysis of voice control frequency databases, the problem of inaccurate voice control commands in virtual characters in multi-person voice environments is solved, and more accurate voice control commands are achieved.

CN120220674APending Publication Date: 2025-06-27HUAIBEI YOUSHENG MEDIA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341317.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing virtual characters based on voice control cannot quickly recognize accurate voice control commands in a multi-person voice environment, resulting in inaccurate voice control commands made by virtual characters, affecting their use.

Method used

By collecting voice control commands from all speakers, timbre identification and feature extraction, timbre signals are generated, and inputting them to the voice control frequency database for similarity matching and frequency analysis, and extracting the tone signal with the highest voice control frequency as voice control command.

Benefits of technology

In a multi-person speaking environment, voice control commands can be accurately identified and extracted, improving the accuracy of voice control commands of virtual characters, and avoiding the problem of inaccurate voice control commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220674A_ABST
    Figure CN120220674A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual character model with a voice control function and a control method thereof, and belongs to the technical field of artificial intelligence, and the control method of the virtual character model with the voice control function comprises the steps: collecting a voice control command; generating a tone signal; inquiring the occurrence frequency of the historical timbre signal; arranging the tone signals matched in the voice control frequency database according to the occurrence frequency, and extracting the tone signal with the highest voice control frequency as a voice control command of the voice control; and generating a voice control instruction. According to the method and the device, the voice control command corresponding to the tone signal is preferentially sent out by adopting the tone signal with the highest matched occurrence frequency, so that an accurate voice control command is generated in a multi-person speaking environment, and the situation that the voice control command cannot be accurately extracted in the multi-person speaking environment, and then the voice control command cannot be generated is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a virtual character model with a voice control function and a control method thereof. Background Art

[0002] Virtual characters generally refer to non-real characters or images created through digital technology. They can exist in various forms, such as 2D animations, 3D models, game characters, virtual idols, etc. In the current era of rapid technological development, the virtual character voice control system is gradually moving from science fiction concepts into real life and becoming a key force driving the transformation of various industries. This system integrates cutting-edge technologies such as speech recognition, natural language processing, and virtual image driving, endowing virtual characters with the ability to "speak" and interact naturally with users, bringing immersive and personalized interaction experiences to users.

[0003] The existing voice control methods for virtual characters mainly include the following: 1. Speech recognition technology: As the "ears" of the system, speech recognition technology is responsible for accurately converting the user's speech into text. From the early template matching algorithms to the current end-to-end models based on deep learning, such as deep neural networks (DNNs), recurrent neural networks (RNNs) and their variants like long short-term memory networks (LSTMs), the recognition accuracy has been greatly improved, and even in noisy environments, it can accurately capture the user's instructions. Taking the speech recognition engine of iFlytek as an example, the recognition accuracy can reach over 98% in a quiet environment, laying a solid foundation for virtual characters to understand the user's intentions. 2. Natural language processing technology: After receiving the speech-to-text result, natural language processing technology comes into play, performing syntactic analysis, semantic understanding, and intention inference on the text. With the help of word vector models (such as Word2Vec, GPT series) and semantic analysis algorithms, the system can understand complex sentence patterns and ambiguous expressions, enabling smooth conversations with users. For example, when the user asks "What's the weather like tomorrow", the system can not only parse the intention of querying the weather but also associate specific time and location information to give an accurate reply. 3. Virtual image driving technology: This is the key to endowing virtual characters with vivid expressiveness. Through speech and lip synchronization algorithms, such as methods based on phoneme and visual feature matching, the virtual character's lip movement is natural and smooth when speaking. At the same time, combined with facial expression generation technology, according to the speech emotion and semantic information, the virtual character is driven to make expressions such as joy, anger, sorrow, and happiness, enhancing emotional interaction. The EchoM imi cV2 tool of Alibaba DAMO Academy has made a breakthrough innovation and can achieve coordinated head and body movements of virtual characters under audio drive, comprehensively improving the expressiveness.

[0004] Currently, in a multi-person voice environment, the virtual characters based on voice control cannot quickly recognize accurate voice control commands, resulting in inaccurate voice control commands made by virtual characters and affecting the use. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a virtual character model with voice control function and its control method.

[0006] The technical solution adopted to solve the above technical problem is: A control method for a virtual character model with voice control function, including:

[0007] Step 1, collect voice control commands of all speakers in this voice control;

[0008] Step 2, conduct timbre discrimination on the collected voice control commands, extract timbre characteristics of different speakers, and generate timbre signals;

[0009] Step 3, input the generated timbre signals into the voice control frequency database, perform similarity matching with the historical timbre signals stored in the voice control frequency database, and query the occurrence frequency of the historical timbre signals;

[0010] Step 4, arrange the timbre signals matched to the voice control frequency database according to the occurrence frequency from high to low, and extract the timbre signal with the highest voice control frequency among them as the voice control command for this voice control;

[0011] Step 5, retrieve the above voice control command and generate a voice control instruction to make the virtual character generate corresponding actions and emit sounds corresponding to the voice control instruction.

[0012] Preferably, the said Step 2 includes the following contents:

[0013] Preprocessing:

[0014] Enhance the high-frequency part through pre-emphasis processing to make the spectrum of the signal tend to be flat;

[0015] Segment the continuous voice signal into short-time frames through frame segmentation processing;

[0016] Reduce the truncation effect at the frame edge through windowing processing;

[0017] Feature extraction:

[0018] Calculate the spectrum of each frame through fast Fourier transform;

[0019] Pass the spectrum through a group of Mel filters through the Mel filter bank;

[0020] Take the logarithm of the output of each filter through logarithmic operation;

[0021] Change the logarithmic energy output by the filter bank through discrete cosine transform to obtain MFCC coefficients;

[0022] Expansion of timbre signal:

[0023] Estimate the linear prediction coefficients of the vocal tract through linear prediction coding;

[0024] Enhance the timbre features through cepstrum lifting.

[0025] Preferably, the voice control frequency database collects voice control commands of different speakers collected during a fixed usage period and generates voice control frequency data.

[0026] Preferably, the method for the voice control frequency library to generate voice control frequency data includes the following:

[0027] Define a fixed usage period as the calculation range of the voice control frequency data;

[0028] Collect voice control commands within each fixed period;

[0029] Process the collected voice signals;

[0030] Extract timbre features from the processed voice signals;

[0031] Conduct frequency analysis on the voice control commands of each speaker to determine the voice control frequency data;

[0032] Generate voice control frequency data according to the frequency analysis results;

[0033] Store the generated voice control frequency data into the database.

[0034] Preferably, the frequency analysis of the voice control commands of each speaker to determine the voice control frequency data includes the following:

[0035] Calculate the fundamental frequency: For each voice frame, calculate its fundamental frequency;

[0036] Frequency distribution statistics: Statistically analyze the fundamental frequencies of all voice commands of each speaker to obtain the frequency distribution.

[0037] Preferably, step 3 includes the following:

[0038] Input the timbre signal generated in the current voice control into the voice control frequency database and verify the integrity of the timbre signal data;

[0039] After inputting into the voice control frequency database, conduct similarity matching;

[0040] For the selected similar historical timbre signals, further query the frequency of occurrence of the historical timbre signals in the voice control frequency database.

[0041] Preferably, after inputting into the voice control frequency database, similarity matching is performed, including the following:

[0042] Compare the characteristic parameters of the input voiceprint signal with the historical voiceprint signals in the voice control frequency database;

[0043] Calculate the similarity between the voiceprint signal input into the voice control frequency database and the historical voiceprint signals through the cosine similarity algorithm.

[0044] Preferably, in step 4, after extracting the voiceprint signal with the highest voice control frequency, first judge the matching accuracy of the extracted voiceprint signal. If the matching accuracy reaches a fixed threshold, then this voiceprint signal is used as the voice control command for this voice control. If the matching accuracy does not reach the fixed threshold, return to step 3.

[0045] A virtual character model with voice control function, including:

[0046] A display module for displaying the virtual character image;

[0047] A voice control frequency database for collecting the voiceprint characteristics of all speakers within a fixed period, generating voiceprint signals from the voiceprint characteristics, and arranging all the voiceprint signals in descending order of the generated frequencies according to the frequencies generated within this period;

[0048] A voice collection module for collecting the voice control commands of all speakers in voice control;

[0049] A voiceprint signal generation module for collecting all the voice control commands, individually discriminating the voiceprints of the voice control commands, extracting the voiceprint characteristics of the speakers corresponding to all the voice control commands, and generating corresponding voiceprint signals based on the voiceprint characteristics;

[0050] A matching module for matching the voiceprint signals to the voice control frequency database, arranging the voiceprint signals in descending order of the voice control frequencies, and extracting the voiceprint signal with the highest voice control frequency as the voice control command for this voice control;

[0051] An instruction module for retrieving the voice control command and generating a voice control instruction to make the virtual character perform corresponding actions and / or emit sounds corresponding to the voice control instruction.

[0052] Preferably, it further includes:

[0053] A command accuracy verification module for detecting whether a voice control command similar to the current voice control command is repeatedly generated within a fixed time after the instruction module ends. If a similar voice control command is re-collected, optimize the matching mechanism of the matching module; otherwise, no optimization is required.

[0054] The beneficial effects of the present invention are as follows:

[0055] 1. In the present invention, by adopting the timbre signal with the highest occurrence frequency to preferentially issue the voice control command corresponding to the timbre signal, accurate voice control instructions can be generated in a multi-person speaking environment, avoiding the inability to accurately extract voice control commands and thus unable to generate voice control instructions in a multi-person speaking environment;

[0056] 2. In the present invention, through step 2 to process the timbre signal, accurate timbre features can be extracted, so as to match the timbre signals in the voice control frequency database;

[0057] 3. In the present invention, by setting a command accuracy verification module, the accuracy of the priority of the current voice control command can be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a schematic flowchart of a control method for a virtual character model with a voice control function in an embodiment of the present invention;

[0059] Figure 2 is a schematic structural diagram of a framework of a virtual character model with a voice control function in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] As Figure 1 - Figure 2 shown, this embodiment provides a virtual character model with a voice control function, including:

[0062] A display module for displaying the virtual character image, and the display module can be a display screen or a device with a display screen;

[0063] A voice control frequency database for collecting the timbre features of all speakers within a fixed period, generating timbre signals based on the timbre features, and arranging all the timbre signals in descending order of the generated frequency according to the frequency of the timbre signals generated within the period. The voice control frequency database collects the voice control commands of different speakers collected during the fixed usage period and generates voice control frequency data. The method for the voice control frequency library to generate voice control frequency data includes the following contents:

[0064] Define a fixed usage period as the calculation range of the voice control frequency data;

[0065] Within each fixed period, collect voice control commands;

[0066] Process the collected voice signals;

[0067] Extract timbre features from the processed voice signals;

[0068] Perform frequency analysis on the voice control commands of each speaker to determine voice control frequency data;

[0069] Generate voice control frequency data according to the frequency analysis results;

[0070] Store the generated voice control frequency data in the database.

[0071] In addition, performing frequency analysis on the voice control commands of each speaker to determine the voice control frequency data includes the following:

[0072] Calculate the fundamental frequency: For each voice frame, calculate its fundamental frequency;

[0073] Frequency distribution statistics: Statistically analyze the fundamental frequencies of all voice commands of each speaker to obtain the frequency distribution;

[0074] A voice acquisition module, used to collect the voice control commands of all speakers in voice control. The voice acquisition module can be a microphone;

[0075] A timbre signal generation module, used to collect all voice control commands, perform timbre discrimination on each voice control command one by one, extract the timbre features of the speakers corresponding to all voice control commands, and generate corresponding timbre signals according to the timbre features;

[0076] A matching module, used to match the timbre signal to the voice control frequency database, arrange the timbre signals according to the high and low voice control frequencies, and extract the timbre signal with the highest voice control frequency as the voice control command for this voice control;

[0077] An instruction module, used to retrieve the voice control command and generate a voice control instruction to make the virtual character generate corresponding actions and / or emit sounds corresponding to the voice control instruction.

[0078] A command accuracy verification module, used to detect whether a voice control command similar to the current voice control command is repeatedly generated within a fixed time after the instruction module ends. If a similar voice control command is re-collected, optimize the matching mechanism of the matching module; otherwise, no optimization is required.

[0079] This embodiment also discloses a voice control method for a virtual character model with voice control function, including:

[0080] Step 1: The voice acquisition module acquires the voice control commands of all speakers in this voice control.

[0081] Step 2: The timbre signal generation module discriminates the timbre of the acquired voice control commands, extracts the timbre features of different speakers, and generates a timbre signal. Specifically, it includes the following contents:

[0082] Preprocessing:

[0083] The high-frequency part is enhanced through pre-emphasis processing to make the spectrum of the signal tend to be flat.

[0084] The continuous voice signal is segmented into short-time frames through frame segmentation processing.

[0085] The truncation effect at the frame edge is reduced through windowing processing.

[0086] Feature extraction:

[0087] The spectrum of each frame is calculated through the fast Fourier transform. The algorithm is as follows:

[0088]

[0089] Where: X(k) is the frequency-domain signal, x(n) is the time-domain signal, N is the number of points of the Fourier transform, and j is the imaginary unit.

[0090] The spectrum is passed through a set of Mel filters through the Mel filter bank. The algorithm is as follows:

[0091]

[0092] Where: M(m) is the output of the mth Mel filter, Hm(k) is the transfer function of the mth Mel filter, and M is the number of filters.

[0093] The logarithm is taken for the output of each filter through logarithmic operation.

[0094] The logarithmic energy output by the filter bank is transformed through the discrete cosine transform to obtain the MFCC coefficients. The discrete cosine transform algorithm is as follows:

[0095]

[0096] Where: C(n) is the nth MFCC coefficient, and N is the number of points of the DCT.

[0097] Timbre signal expansion:

[0098] The linear prediction coefficients of the vocal tract are estimated through linear prediction coding.

[0099] The timbre features are enhanced through cepstrum lifting.

[0100] In addition, the algorithm for converting timbre features into timbre signals is as follows:

[0101]

[0102] Where: S(f) is the generated timbre signal, C(n) is the nth timbre feature coefficient, g(f) is a window function used to smooth the timbre signal, and F is the frequency range of the timbre signal.

[0103] Step 3: The matching module inputs the generated timbre signal into the voice control frequency database and performs similarity matching with the historical timbre signals stored in the voice control frequency database to query the occurrence frequency of the historical timbre signal. Specifically, it includes the following contents:

[0104] Input the timbre signal generated in this voice control into the voice control frequency database and verify the integrity of the timbre signal data;

[0105] After inputting into the voice control frequency database, perform similarity matching, which specifically includes the following contents:

[0106] Compare the characteristic parameters of the input timbre signal with the historical timbre signals in the voice control frequency database;

[0107] Calculate the similarity between the timbre signal input into the voice control frequency database and the historical timbre signal through the cosine similarity algorithm;

[0108] For the selected similar historical timbre signals, further query the frequency of occurrence of the historical timbre signal in the voice control frequency database.

[0109] Step 4: The matching module arranges the timbre signals matched to the voice control frequency database according to the frequency of occurrence from high to low, extracts the timbre signal with the highest voice control frequency among them, and then judges the matching accuracy of the extracted timbre signal. If the matching accuracy reaches a fixed threshold, the timbre signal is used as the voice control command for this voice control. If the matching accuracy does not reach the fixed threshold, return to Step 3;

[0110] Step 5: The instruction module retrieves the above voice control command and generates a voice control instruction to make the virtual character generate corresponding actions and emit sounds corresponding to the voice control instruction.

[0111] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A method for controlling a virtual character model with a voice control function, characterized in that: include: Step 1: Collect the voice control commands of all speakers in this voice control; Step 2: Perform timbre identification on the collected voice control commands, extract timbre features of different speakers, and generate timbre signals; Step 3, input the generated timbre signal into the voice control frequency database, and perform similarity matching with the historical timbre signal stored in the voice control frequency database to query the occurrence frequency of the historical timbre signal; Step 4: Arrange the timbre signals matched to the voice control frequency database according to the frequency of occurrence, and extract the timbre signal with the highest voice control frequency as the voice control command for this voice control; Step 5: Retrieve the above voice control command and generate a voice control instruction, so that the virtual character performs corresponding actions and makes sounds corresponding to the voice control instruction.

2. The control method of a virtual character model with voice control function according to claim 1, characterized in that: The step 2 includes the following contents: Preprocessing: The high frequency part is enhanced by pre-emphasis processing to make the spectrum of the signal flat; The continuous speech signal is divided into short time frames through frame processing; Reduce the truncation effect at the frame edge by windowing; Feature extraction: The spectrum of each frame is calculated by fast Fourier transform; The spectrum is passed through a set of Mel filters through a Mel filter bank; Take the logarithm of each filter output by logarithmic operation; The logarithmic energy of the filter bank output is changed by discrete cosine transform to obtain MFCC coefficients; Tone signal expansion: By linear prediction coding, the linear prediction coefficients of the vocal channel are estimated; Enhance the timbre characteristics through cepstral enhancement.

3. The control method of a virtual character model with voice control function according to claim 1, characterized in that: The voice control frequency database collects voice control commands from different speakers in a fixed usage period and generates voice control frequency data.

4. The control method of the virtual character model with voice control function according to claim 3, characterized in that: The method for generating voice control frequency data from the voice control frequency library includes the following contents: Define a fixed usage period as the calculation range of voice control frequency data; In each fixed period, voice control commands are collected; Processing the collected voice signals; Extracting timbre features from the processed speech signal; performing frequency analysis on each speaker's voice control commands to determine voice control frequency data; Generate voice control frequency data according to the frequency analysis result; The generated voice control frequency data is stored in a database.

5. The control method of the virtual character model with voice control function according to claim 4, characterized in that: The frequency analysis of the voice control command of each speaker to determine the voice control frequency data includes the following contents: Calculate the fundamental frequency: For each speech frame, calculate its fundamental frequency; Frequency distribution statistics: The fundamental frequencies of all voice commands of each speaker are counted to obtain the frequency distribution.

6. The control method of a virtual character model with voice control function according to claim 1, characterized in that: The step 3 includes the following contents: Input the timbre signal generated in this voice control into the voice control frequency database, and verify the integrity of the timbre signal data; After inputting the voice control frequency database, similarity matching is performed; For the screened similar historical timbre signals, the frequency of occurrence of the historical timbre signals in the voice control frequency database is further queried.

7. The control method of the virtual character model with voice control function according to claim 6, characterized in that: After the voice control frequency database is input, similarity matching is performed, including the following: comparing the input timbre signal with characteristic parameters of historical timbre signals in a voice control frequency database; The similarity between the timbre signal input into the voice control frequency database and the historical timbre signal is calculated by the cosine similarity algorithm.

8. The control method of a virtual character model with voice control function according to claim 1, characterized in that: In step 4, after extracting the timbre signal with the highest voice control frequency, the extracted timbre signal is first judged for matching accuracy. If the matching accuracy reaches a fixed threshold, the timbre signal is used as the voice control command for this voice control. If the matching accuracy does not reach the fixed threshold, return to step 3.

9. A virtual character model with voice control function, characterized in that: include: A display module, used for displaying a virtual character image; The voice control frequency database is used to collect the timbre characteristics of all speakers within a fixed period, and generate timbre signals from the timbre characteristics, and then arrange all timbre signals according to the frequency of timbre signals generated within the period; A voice collection module, used to collect voice control commands of all speakers in voice control; The timbre signal generation module is used to collect all voice control commands, identify the timbre of the voice control commands one by one, extract the timbre features of the speakers corresponding to all voice control commands, and generate corresponding timbre signals based on the timbre features; A matching module is used to match the timbre signal to the voice control frequency database, and arrange the timbre signals according to the high and low voice control frequencies, and extract the timbre signal with the highest voice control frequency as the voice control command for this voice control; The command module is used to retrieve the voice control command and generate the voice control command to make the virtual character perform corresponding actions and / or emit sounds corresponding to the voice control command.

10. The virtual character model with voice control function according to claim 9, characterized in that: Also includes: The command accuracy verification module is used to detect whether a voice control command similar to the current voice control command is repeatedly generated within a fixed time after the instruction module ends. If a similar voice control command is re-collected, the matching mechanism of the matching module is optimized, otherwise no optimization is required.