Voice interaction method and system of mobile digital human

Through the anti-neural network and digital human emotion analysis model, the speech waveform capture is optimized, and the speech signal fluctuation feature group is generated, which solves the speech recognition and dialogue continuity problems of mobile digital humans in non-fixed scenarios, and achieves efficient and natural voice interaction effects.

CN120279896AInactive Publication Date: 2025-07-08XIAN WUKONG INTELLIGENT TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510643104.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, mobile digital people have low speech recognition accuracy in non-fixed scenarios, poor dialogue continuity, lack of contextual correlation of response content, and lack of dynamic adjustment ability of voice feedback, resulting in poor interaction effect.

Method used

Data is input through the voice acquisition channel, the sound wave signal timing is extracted, combined with the adversarial neural network and digital human emotion analysis model, a speech signal fluctuation feature group is generated, mutation candidate characteristics are screened, weighted coefficient processing and response index establishment are carried out, and dynamic adjustment of speech output and dynamic expression synchronization of virtual characters are optimized.

Benefits of technology

It improves the coherence of voice interaction, the semantic accuracy of response and the user immersive interactive experience, enhances the expressiveness and naturalness of voice output, and optimizes the response efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279896A_ABST
    Figure CN120279896A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction, in particular to a voice interaction method and system for a mobile digital human, and the method comprises the steps: inputting data through a voice collection channel, reading a sound wave signal time sequence, extracting a waveform amplitude value, and recording a timestamp, and synchronously extracting and coding time domain and frequency domain changes implied in a conventional voice input signal. The method improves the capturing precision of tiny mutation characteristics in voice signals, introduces a generation and discrimination mechanism of an adversarial neural network through the generation and screening process of a mutation candidate characteristic group, enables a generator to dynamically generate diversified amplitude and frequency change modes, further improves the expressive force and naturalness of voice output, and improves the voice recognition accuracy. By performing weighting coefficient addition processing on the product of the amplitude and the frequency change, the voice output intensity can be reflected more accurately, the response performance of the voice is optimized, and the response efficiency and the stability of the system are further improved by setting the bidirectional mapping and the freezing state flag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular to a voice interaction method and system for a mobile digital human. Background Art

[0002] Voice interaction technology aims to exchange human-computer information based on voice signals, including modules such as voice recognition, semantic understanding, speech synthesis and voice drive. It is widely used in scenarios such as smart terminals, service robots, and human-computer dialogue systems. It is committed to achieving efficient command transmission, information feedback and task control through natural language, and plays an important role in improving the convenience of interaction and the intelligence of response.

[0003] A method for voice interaction of mobile digital human, built on artificial intelligence technology, specifically integrates multiple artificial intelligence subsystems such as speech recognition, semantic understanding, language generation and behavior control, to enable digital human to have natural language interaction capabilities in mobile scenarios, aiming to solve the technical problems of low speech recognition accuracy, poor dialogue continuity, lack of contextual relevance of response content, etc. in existing technologies of mobile digital human in non-fixed scenarios. By constructing an interactive process with perception, understanding, generation and feedback capabilities, digital human can conduct multiple rounds of contextual voice dialogues according to environmental changes and user needs while on the move, thereby improving the coherence of voice interaction, semantic accuracy of response and immersive interaction experience of users.

[0004] Although traditional methods can achieve basic speech recognition and response generation in voice interaction, they still have obvious deficiencies in speech diversity and expressiveness. Existing voice feedback is usually generated based on preset rules and templates, and lacks the ability to flexibly adjust according to the dynamic changes of interaction scenarios. Static voice output methods often cannot fully adapt to the diversity of user needs, resulting in users feeling mechanical feedback in the interaction, reducing the naturalness and fun of the interaction. The style of voice output in traditional methods lacks dynamic adjustment capabilities, and cannot adjust the emotional color and tone of voice according to different situations, which makes the voice interaction system limited in situational adaptability. Existing technologies have not been able to effectively achieve fine adjustment and optimization of amplitude changes and frequency differences, resulting in poor performance of voice feedback in some complex scenarios, affecting the smoothness of the interaction effect and user satisfaction. Existing technologies have failed to achieve an ideal optimization level in terms of voice output diversity, situational adaptability and interaction tension, and cannot meet the interaction needs of users in highly personalized and changeable scenarios. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a voice interaction method and system for a mobile digital human.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A voice interaction method for a mobile digital human, comprising the following steps:

[0007] Step 1: Based on the input data of the voice acquisition channel, read the time sequence of the acoustic wave signal. By sequentially extracting the waveform amplitude values of the sampling points and recording the corresponding timestamps, extracting the frequency values and performing subtraction processing on adjacent frequency values, and combining the natural language processing algorithm in the digital human voice synthesis technology, further optimize the capture accuracy and analysis process of the voice waveform, and generate a voice signal fluctuation feature group;

[0008] Step 2: Based on the voice signal fluctuation feature group, screen the amplitude and frequency data within the boundary range, generate a mutation candidate feature group, use an adversarial neural network to expand and generate and discriminate effective pairs, obtain an effective mutation pair group, assign numbers to screen mutation segments, and combine the digital human emotion analysis model to enhance the emotion-level discrimination and voice matching of voice mutation features, and generate a mutation voice segment identification set;

[0009] Step 3: Based on the mutation voice segment identification set, extract the frequency values and amplitude values, classify and organize the data within each segment according to the channel number, obtain the product of the amplitude change value and the frequency difference, sequentially perform weighted coefficient addition processing and record the amplitude response and frequency response values, and at the same time, by introducing the digital human dynamic expression and posture synchronization algorithm, generate a segment intensity analysis result;

[0010] Step 4: Based on the segment intensity analysis result, obtain the response numerical values corresponding to the channel numbers through bidirectional mapping, establish an index relationship between the channel numbers and the response intensities, set a freeze state flag and write it into the storage buffer structure, and combine the digital human real-time voice synthesis optimization technology to generate a voice output freeze buffer structure;

[0011] Step 5: Based on the voice output freeze buffer structure, obtain the difference between the current value and the historical record value of the output vector, screen the difference and locate the vector index numbers that meet the threshold criteria, arrange and assemble the index combinations into a matrix structure. In this process, through the digital human voice-visual interaction algorithm, realize the natural matching of voice output and virtual character actions, and generate a voice response feature output matrix.

[0012] As a further solution of the present invention, the specific steps for generating the voice signal fluctuation feature group are as follows:

[0013] Based on the input data of the voice acquisition channel, read the time sequence of the acoustic wave signal, extract the acoustic wave amplitude value of the sampling point, use the amplitude value direct assignment method to record the corresponding timestamp of the sampling point, and combine the audio feature processing method in the digital human voice synthesis technology. By optimizing the sampling data processing, improve the capture accuracy of the voice waveform, combine the sampling point amplitude and timestamp in chronological order, establish a unified sequence structure, and form a waveform amplitude time sequence data group;

[0014] Based on the waveform amplitude time series data set, calculate the amplitude change rate in adjacent time periods, extract the change rate by dividing the difference by the time difference, combine with the digital human emotion analysis model, further analyze the emotion level in the abnormal change segment, and optimize the voice matching and processing of abnormal fluctuations in combination with the emotion discrimination result. Screen the continuously changing amplitude values, set a threshold range to discriminate the abnormal change rate section, extract it into an abnormal segment, and obtain the amplitude change feature group;

[0015] Based on the amplitude change feature group, extract the main frequency in the corresponding time period, subtract adjacent frequency values in turn by the frequency sampling value difference method, combine with the digital human voice-visual interaction algorithm, synchronize the frequency fluctuation data with the dynamic expressions and postures of the digital human, optimize the collaborative performance between the voice signal and the virtual character's actions, combine the frequency differences to form a continuous change sequence, screen the complete cycle frequency fluctuation change data, and generate the voice signal fluctuation feature group.

[0016] As a further solution of the present invention, the specific steps for generating the mutant voice segment identification set are as follows:

[0017] Based on the voice signal fluctuation feature group, extract the amplitude change value and the frequency difference, sort out the amplitude change value and the frequency difference data, locate the amplitude change value and the frequency difference, set the limit extreme values of the amplitude change value and the frequency difference, combine with the digital human emotion analysis model, analyze the emotion characteristics in the voice fluctuation, so as to enhance the emotion level discrimination of the voice mutation characteristics; compare the amplitude change value and the frequency difference, screen out the data pairs that exceed the limit, establish a data set that meets the limit conditions, and generate a mutant candidate feature group;

[0018] Based on the mutant candidate feature group, use the adversarial neural network generator to generate and expand the combined samples of the amplitude change value and the frequency difference, discriminate the authenticity of the generated samples and the original samples, screen out the non-genuine mutant samples, retain the paired units of the amplitude change value and the frequency difference that meet the mutant feature criteria, combine with the digital human voice-visual interaction algorithm, optimize the synchronization of the generated mutant samples with the actions and expressions of the virtual character, so as to ensure the coordination between the voice mutation and the emotional expression of the virtual character, combine the effective unit data groups, and obtain the effective mutant pairing group;

[0019] Based on the effective mutant pairing group, assign a sampling segment number, use the corresponding index method of the amplitude change amplitude and the frequency change amplitude to mark the data number, screen the index numbers that match the mutant amplitude and the mutant frequency, combine them to form a continuous numbered paragraph, combine with the digital human real-time voice synthesis optimization technology, synchronize the numbered paragraph with the dynamic expressions and postures of the virtual character, improve the interactive effect of the voice and the action, and then sort out the numbered set to generate the mutant voice segment identification set.

[0020] As a further solution of the present invention, the adversarial neural network is calculated according to the formula:

[0021]

[0022] Where: is the loss function, G is the generator neural network module, which receives the input noise vector to generate mutant samples, D is the discriminator neural network module, E is the expectation operator, x is the original real sample, p data (x) is the data distribution of the real sample, log is the logarithm operator, D(x) is the discriminant output value of the discriminator for the input real sample x, z is the noise vector, p z (z) is the sampling distribution of the noise variable z, D(G(z)) is the discriminant output value of the discriminator for the sample G(z) generated by the generator, 1 - D(G(z)) is the probability that the generated sample is judged as a false sample, λ is the weight adjustment coefficient, x' is the mutant sample generated by the generator, p mut (x') is the sampling distribution of the generated mutant sample x', x target is the target mutation standard sample, ||x' - x target || is the Euclidean distance between the generated mutant sample x' and the target sample x target ||x' - x target || 2 is the square value of the above Euclidean distance.

[0023] As a further solution of the present invention, the amplitude change value and frequency difference data refer to the amplitude change value and frequency difference data extracted first, arranging the amplitude change value and frequency difference index data, sorting the amplitude change values in ascending order, synchronously adjusting the corresponding frequency difference index positions with the index positions after sorting the amplitude change values, normalizing the frequency difference data, scaling the frequency difference data to a set interval, establishing a two-way index table of the amplitude change value and the frequency difference normalization result, batch-merging the amplitude change value intervals, and completing the synchronous collation of the amplitude change value and frequency difference data.

[0024] As a further solution of the present invention, the specific steps for generating the segment intensity analysis result are as follows:

[0025] Based on the mutant speech segment identification set, extract the frequency values and amplitude values within the segment, read and extract the frequency change values and amplitude change values of the sampling points within the corresponding segment one by one, and combine with the digital human dynamic expression and posture synchronization algorithm to synchronize these changes with the actions and expressions of the virtual character while extracting the frequency and amplitude changes, enhancing the coordination between the voice and visual effects, classifying the channel numbers to which the sampling points belong, classifying the channel numbers of the data content, establishing the corresponding relationship between the channel numbers and the data content, and generating a channel classification data list;

[0026] Based on the channel classification data list, calculate the product of the amplitude change value and the frequency change value for each channel number. Match the amplitude change values item by item and multiply them with the corresponding frequency change values. Arrange the segment numbers of the product values in order. Combining with the digital human speech synthesis technology, while calculating the product values, optimize the speech waveform. Group and combine them into data sequences according to the channel numbers, integrate the data groups corresponding to the channel numbers, and obtain the product combination sequence;

[0027] Based on the product combination sequence, set and assign weighting coefficients, multiply each product value within each group with the weighting coefficient item by item, accumulate the product of the data and the weighting coefficient, and record the amplitude response value and the frequency response value. Combining with the digital human real-time speech synthesis optimization technology, optimize the speech output during the weighting process, arrange the response results in the order of segment numbers, combine the response data corresponding to all segments, and generate the segment intensity analysis result.

[0028] As a further solution of the present invention, the specific steps for generating the speech output freeze cache structure are as follows:

[0029] Based on the segment intensity analysis result, extract the channel number and the response value, locate the channel number and the corresponding response value in the indexing order positioning method, and synchronously arrange the channel number list and the response value list. Combining with the digital human dynamic expression and posture synchronization algorithm, optimize the dynamic expression and posture of the virtual character while synchronizing the channel number and the response value. Match the channel number and the response value according to the number mapping method, establish a two-way access index relationship, and generate a channel response two-way mapping table;

[0030] Based on the channel response two-way mapping table, establish the corresponding relationship between the channel number and the response intensity, rearrange the indexing order in the ascending order of the channel number, refer to the response value in the mapping table to assign the number node for the response intensity, combine with the digital human speech synthesis technology, optimize the response intensity, organize it into a unified structure according to the number set method, combine the content of each number and the corresponding response intensity, and obtain the channel response index set;

[0031] Based on the channel response index set, set the freeze status flag, compare the response intensity corresponding to the channel number using the threshold screening method to obtain the value, screen out the channel numbers that meet the freeze conditions, set the freeze flag to identify the freeze status, combine with the digital human real-time speech synthesis optimization technology, enhance the accuracy and stability of the speech output in the management of the freeze status, obtain the channel numbers in order and write them into the storage buffer unit in sequence, combine each freeze status data, and generate the speech output freeze cache structure.

[0032] As a further solution of the present invention, the channel number and the corresponding response value of the index sequential positioning method are specifically as follows: extract the channel number and the response value, index the original arrangement order of the channel numbers, map the channel numbers and the corresponding response values, retrieve the response values according to the index order of the channel numbers, locate the position of the actual response value corresponding to the channel number, and establish a synchronous relationship between the channel number index pointer and the response value index pointer to complete the mapping binding from the channel number index to the response value index.

[0033] As a further solution of the present invention, the specific steps for generating the voice response feature output matrix are as follows:

[0034] Based on the voice output freeze cache structure, obtain the output vector values, sequentially traverse to locate the output values corresponding to the channel numbers, and retrieve the historical record values corresponding to the channel numbers. Combining the digital human voice synthesis optimization technology, optimize the capture accuracy of the voice data when obtaining the output vector, subtract the current output value from the historical record value item by item, calculate the channel number difference, combine the difference data, and generate an output vector difference set;

[0035] Based on the output vector difference set, screen and process the differences, extract the absolute values of the differences item by item, and then compare the absolute values with the set threshold item by item to screen out the channel numbers whose absolute values of the differences meet the preset values. Combining the digital human real-time voice synthesis technology, optimize the difference screening process, extract the corresponding index number set, sort the screened channel numbers in ascending order, and obtain the screened index number group;

[0036] Based on the screened index number group, arrange the index numbers, organize the index numbers of each channel, arrange the index numbers row by row according to the specified row-column matrix rule. Combining the digital human dynamic expression and posture synchronization algorithm, synchronize the voice output with the dynamic expressions and postures of the virtual characters during the index number sorting process, combine continuous index segments to form row vectors, and stack the row vectors to form the main body of the matrix, and output the overall data structure of the matrix to generate the voice response feature output matrix.

[0037] A voice interaction system for a mobile digital human, the voice interaction system for the mobile digital human is used to execute the above-mentioned voice interaction method for the mobile digital human, and the system includes:

[0038] Sound wave feature extraction module: Based on the sound wave input channel data, read the time series content of the sound wave signal, extract the waveform amplitude value of each sampling point, record the corresponding timestamp data, combine the digital human voice synthesis technology, optimize the sampling data through frequency domain transformation processing, improve the accuracy of voice waveform capture, extract the frequency value of each time point through frequency domain transformation processing, calculate the difference between the frequency values of adjacent time points, generate two data sets of amplitude change sequence and frequency difference sequence, and combine the amplitude change sequence and the frequency difference sequence to obtain the voice signal fluctuation feature group;

[0039] Mutation Feature Screening Module: Based on the speech signal fluctuation feature group, screen the amplitude change values within the preset amplitude limit range in the amplitude change sequence. At the same time, use an adversarial neural network to screen the frequency difference values within the preset frequency limit range in the frequency difference sequence. Combine with the digital human emotion analysis model to further analyze the emotional fluctuations in the mutation features. Combine the screened amplitude change values and frequency differences into a paired unit, discriminate and classify the generated results and the real data, eliminate the pseudo-mutation samples with a large deviation from the real distribution, retain the set of effective paired units that meet the mutation feature criteria, process the set of effective paired units and number them, screen the mutation segment indexes, integrate the set of mutation segment indexes, and generate a mutation speech segment identification set;

[0040] Fragment Intensity Analysis Module: Based on the mutation speech segment identification set, extract the frequency values and amplitude values within each segment, classify and sort the frequency values and amplitude values of the acoustic wave input channel numbers. Combine with the digital human speech-visual interaction algorithm to optimize the synchronization of the speech and the dynamic expressions of the virtual character during the analysis process. Calculate the product result of the amplitude change value and the frequency difference within each channel, perform the weighted coefficient addition process on the product result, record the total amplitude response value and the total frequency response value within each channel, combine the total amplitude response value and the total frequency response value, establish a response intensity set corresponding to the segment, and generate a fragment intensity analysis result;

[0041] Response Index Establishment Module: Based on the fragment intensity analysis result, establish a two-way mapping relationship between the channel number and the response intensity, extract the total amplitude response value and the total frequency response value data corresponding to each channel number, index and sort according to the number order to establish a channel index table, set a freeze state flag, and perform a write process to the storage buffer for the freeze flag value; Combine with the digital human real-time speech synthesis optimization technology to optimize the processing and storage of the response data to ensure the fluency of the voice output and the synchronization of the virtual character's actions; Establish a freeze state buffer table structure and generate a voice output freeze cache structure;

[0042] Feature Output Generation Module: Based on the voice output freeze cache structure, extract the output vector value under the current time segment and the output vector value recorded under the corresponding historical time segment, and calculate the difference between the current value and the historical value; Combine with the digital human speech synthesis optimization technology to enhance the smoothness of the output difference, screen the data points in the difference set whose absolute value is greater than the threshold standard, extract the index numbers of the data points that meet the conditions, arrange them in combination according to the index numbers, assemble them into a standard matrix structure, establish a matrix index of the acoustic wave channel number, and generate a voice response feature output matrix.

[0043] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0044] 1. In the present invention, by inputting data through the voice acquisition channel to read the time sequence of the acoustic wave signal, extracting the waveform amplitude value and recording the timestamp, it is possible to synchronously extract and encode the time-domain and frequency-domain changes hidden in the traditional voice input signal, increasing the capture accuracy of the minute mutation characteristics in the voice signal;

[0045] 2. In the present invention, through the generation and screening process of the mutation candidate feature group, the generation and discrimination mechanisms of the adversarial neural network are introduced, enabling the generator to dynamically generate diverse amplitude and frequency change patterns, and effectively judging the authenticity through the discriminator, further enhancing the expressiveness and naturalness of the voice output;

[0046] 3. In the present invention, by performing weighted coefficient addition processing on the product of the amplitude and frequency changes, it is possible to more accurately reflect the intensity of the voice output, optimizing the response performance of the voice. In the final stage of response generation, the setting of the bidirectional mapping and the frozen state flag further improves the response efficiency and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic diagram of the working process of the present invention;

[0048] Figure 2 is a system flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] Embodiment 1: Please refer to Figure 1 , the present invention provides a technical solution: a voice interaction method for a mobile digital human, including the following steps:

[0051] Step 1: Based on the input data through the voice acquisition channel, read the time sequence of the acoustic wave signal. By sequentially extracting the waveform amplitude value for each sampling point and recording the corresponding timestamp, extracting the frequency value and performing subtraction processing on adjacent frequency values, and further optimizing the capture accuracy and analysis process of the voice waveform by combining the natural language processing algorithm in the digital human voice synthesis technology, a voice signal fluctuation feature group is generated;

[0052] Step 2: Based on the voice signal fluctuation feature group, screen the amplitude and frequency data within the boundary range to generate a mutation candidate feature group. Use the adversarial neural network to expand and generate and discriminate valid pairs, obtain a valid mutation pair group, assign numbers to screen mutation segments, and combine the digital human emotion analysis model to enhance the emotion-level discrimination and voice matching of the voice mutation characteristics, generating a mutation voice segment identification set;

[0053] Step 3: Based on the set of mutated voice segment identifiers, extract the frequency values and amplitude values, classify and organize the data within each segment according to the channel number, obtain the product of the amplitude change value and the frequency difference, sequentially perform weighted coefficient addition processing and record the amplitude response and frequency response values. At the same time, by introducing a digital human dynamic expression and posture synchronization algorithm, generate the segment intensity analysis result;

[0054] Step 4: Based on the segment intensity analysis result, obtain the response values corresponding to the channel numbers through bidirectional mapping, establish the index relationship between the channel numbers and the response intensity, set the freeze state flag and write it into the storage buffer structure, and combine the digital human real-time speech synthesis optimization technology to generate the voice output freeze cache structure;

[0055] Step 5: Based on the voice output freeze cache structure, obtain the difference between the current value and the historical record value of the output vector, filter the difference and locate the vector index number that meets the threshold standard, arrange and assemble the index combinations into a matrix structure. In this process, through the digital human speech-visual interaction algorithm, realize the natural matching of the voice output and the virtual character's actions, and generate the voice response feature output matrix.

[0056] The specific steps to generate the voice signal fluctuation feature group are as follows:

[0057] Based on the input data of the voice acquisition channel, read the time sequence of the sound wave signal, extract the sound wave amplitude value of the sampling point, record the time stamp corresponding to the sampling point by using the direct assignment method of the amplitude value, and combine the audio feature processing method in the digital human speech synthesis technology. By optimizing the sampling data processing, improve the accuracy of voice waveform capture, combine the amplitude and time stamp of the sampling point in chronological order, and establish a unified sequence structure to form a waveform amplitude time sequence data group;

[0058] Based on the waveform amplitude time sequence data group, calculate the amplitude change rate in adjacent time periods, extract the change rate by using the method of dividing the difference by the time difference, combine the digital human emotion analysis model, further analyze the emotion level in the abnormal change segment, and optimize the voice matching and processing of the abnormal fluctuation in combination with the emotion discrimination result. Screen the continuously changing amplitude values, set the threshold range to discriminate the abnormal change rate section, and extract it into abnormal segments to obtain the amplitude change feature group;

[0059] Based on the amplitude change feature group, extract the main frequency within the corresponding time period, subtract the adjacent frequency values sequentially by using the frequency sampling value difference method, combine the digital human speech-visual interaction algorithm, synchronize the frequency fluctuation data with the dynamic expression and posture of the digital character, optimize the collaborative performance between the voice signal and the virtual character's actions, combine the frequency differences to form a continuously changing sequence, and screen the complete cycle frequency fluctuation change data to generate the voice signal fluctuation feature group;

[0060] Based on the input data of the voice acquisition channel, read the time sequence of the acoustic wave signal, extract the amplitude value of the acoustic wave at the sampling point, record the amplitude value corresponding to the time stamp of the sampling point in this way, use the DataFrame object in the Pandas library of Python to store the amplitude value and time stamp of the sampling point as data items, combine the amplitude and time stamp of the sampling point in chronological order, use the sort_values command to sort the time stamp column, establish a sequence structure, generate a waveform amplitude time sequence data group, use the read_csv method of the Pandas library to read the acoustic wave signal data from the data stream of the voice acquisition device, extract the amplitude value and time stamp column of the sampling point, construct a table data structure containing data items, apply the sort_values function to sort the time stamp in ascending order, obtain a set of sampling data arranged in chronological order, and form a structured waveform amplitude time sequence data group;

[0061] Based on the waveform amplitude time sequence data group, calculate the amplitude change rate in adjacent time periods, extract the change rate of the difference divided by the time difference, use the diff function in the NumPy library of Python to calculate the amplitude difference between consecutive time points, divide the amplitude difference by the corresponding time difference to obtain the amplitude change rate, use numpy.diff to calculate the amplitude difference between adjacent sampling points, calculate the time difference using the time stamp difference, screen the continuously changing amplitude values, set a threshold range to identify abnormal change rate sections, use the apply function of Pandas to filter out fluctuations with a rate lower than the set threshold, through the set threshold, apply the DataFrame.query method to filter out amplitude changes with a change rate exceeding the set threshold, extract them into abnormal segments, and generate an amplitude change feature group;

[0062] Based on the amplitude change feature group, extract the main frequency within the corresponding time period, subtract adjacent frequency values successively using the frequency sampling value difference method, use the find_peaks method in scipy.signal to perform peak detection on the frequency values, extract the main frequency value of the frequency signal, calculate the difference between adjacent frequency values, use the NumPy.diff method to subtract the frequency values pairwise, generate a frequency difference sequence, combine the frequency differences to form a continuously changing sequence, and through the groupby and agg functions of Pandas, perform periodic screening on the frequency difference data, screen the frequency fluctuation change data of the complete cycle, and generate a voice signal fluctuation feature group.

[0063] The specific steps to generate a set of identification of mutant voice segments are as follows:

[0064] Based on the voice signal fluctuation feature group, extract the amplitude change value and the frequency difference value, organize the amplitude change value and the frequency difference value data, locate the amplitude change value and the frequency difference value, set the limit extreme values of the amplitude change value and the frequency difference value, and combine with the digital human emotion analysis model to analyze the emotion features in the voice fluctuation, so as to enhance the emotion level discrimination of the voice mutation features; compare the amplitude change value and the frequency difference value, screen out the data pairs that exceed the limit, establish a data set that meets the limit conditions, and generate a mutation candidate feature group;

[0065] Based on the mutation candidate feature group, use the generative adversarial network generator to generate and expand the combined samples of the amplitude change value and the frequency difference value, discriminate the authenticity of the generated samples and the original samples, screen out the non-genuine mutation samples, retain the paired units of the amplitude change value and the frequency difference value that meet the mutation feature criteria, and combine with the digital human voice-visual interaction algorithm to optimize the synchronization of the generated mutation samples with the actions and expressions of the virtual character, so as to ensure the coordination of the voice mutation and the emotional expression of the virtual character, combine the effective unit data groups, and obtain the effective mutation pairing groups;

[0066] Based on the effective mutation pairing groups, assign the sampling segment numbers, use the corresponding index method of the amplitude change amplitude and the frequency change amplitude to mark the data numbers, screen out the index numbers that match the mutation amplitude and the mutation frequency, combine them to form a continuous numbered paragraph, and combine with the digital human real-time voice synthesis optimization technology to synchronize the numbered paragraph with the dynamic expressions and postures of the virtual character, improve the interaction effect of the voice and the action, and then organize the numbered set to generate a mutation voice segment identification set;

[0067] Based on the voice signal fluctuation feature group, use the data screening and limit comparison algorithm to extract the amplitude change value and the frequency difference value. Use the apply method in the DataFrame structure in the Pandas library in Python to process the amplitude change value and the frequency difference value. Set the limit extreme value parameter min_threshold to -0.3 and max_threshold to 0.3. Apply the conditional screening expression df.query('min_threshold <= value <= max_threshold') for range screening, screen out the data pairs that exceed the limit, organize the screened amplitude change value and frequency difference value data, establish a data set that meets the limit conditions, and generate a mutation candidate feature group;

[0068] Based on the mutation candidate feature group, an adversarial neural network is used to generate a discrimination algorithm to expand the combined samples of the amplitude change value and the frequency difference value. Using the PyTorch framework, a generator network G is defined. The input parameters include the initial amplitude change value u, the initial frequency difference value v, the amplitude change rate s, and the frequency change gradient t. A four-dimensional vector input method is used for sample generation. The cross-entropy loss function BinaryCrossEntropyLoss is used to calculate the classification loss between the generated sample G(u, v, s, t) and the real samples a, b. The weights of the generator and discriminator are optimized through backpropagation, the samples judged to be forged are screened out, the paired units of the amplitude change value and the frequency difference judged to be real are combined, and an effective mutation pairing group is obtained;

[0069] Based on the effective mutation pairing group, a sequence index and paragraph numbering algorithm is used to mark the amplitude change amplitude and the frequency change amplitude. The arange method in NumPy is used to generate a numbered index array, and numbers are assigned synchronously according to the mutation amplitude change position and the frequency change position. Boolean indexing is used to filter the index numbers where the mutation amplitude and the mutation frequency match. The groupby method of Pandas is used to group by consecutive numbers to form consecutive numbered paragraphs, and then the paragraph numbers are sorted to generate a mutation speech segment identification set.

[0070] The adversarial neural network, according to the formula:

[0071]

[0072] Where: is the loss function, G is the generator neural network module, which receives the input noise vector to generate mutation samples, D is the discriminator neural network module, E is the expectation operator, x is the original real sample, ~ is the sampling symbol from the distribution, p data (x) is the data distribution of the real sample, log is the logarithmic operator, D(x) is the discriminator's discrimination output value for the input real sample x, z is the noise vector, p z (z) is the sampling distribution of the noise variable z, D(G(z)) is the discriminator's discrimination output value for the sample G(z) generated by the generator, 1 - D(G(z)) is the probability that the generated sample is judged to be a fake sample, λ is the weight adjustment coefficient, x′ is the mutation sample generated by the generator, p mut (x′) is the sampling distribution of the generated mutation sample x′, x target is the target mutation standard sample, ||x′ - x target || is the Euclidean distance between the generated mutation sample x′ and the target sample x target between, ||x′ - x target || 2 is the square value of the above Euclidean distance.

[0073] Execution process: In a voice interaction method and system for mobile digital humans, a noise vector z is extracted from the noise distribution p z (z) by a noise sampling module, and the noise vector z is input into a generator G. The neural network processes to generate a mutated sample x'. The generated sample G(z) is input into a discriminator D, and the discriminator D outputs the authenticity score D(G(z)) of the generated sample. The original real sample x is sampled from the real data distribution p data (x) and then input into the discriminator D, which outputs the authenticity score D(x) of the real sample. Based on the basic loss function of the generative adversarial network and the real sample score and the generated sample score are calculated respectively, and the two scores are accumulated to obtain the loss of the adversarial part. An additional mutation standard constraint is introduced, and the generated mutated sample x' is compared with a preset target mutation standard sample x target in terms of features, and the Euclidean distance square ||x' - x target || target between them is calculated to measure the compliance degree of the generated sample x' with the standard. The expected value operation E 2 is performed on the mutated sample distribution p mut (x') to obtain E x′~pmut(x′) [||x' - x target || 2 , the expected value of compliance is obtained. The adjusted compliance term is added to the total loss by setting a weight adjustment coefficient λ to form an improved loss function The loss function is minimized to optimize the training processes of the generator G and the discriminator D, and effective amplitude change value and frequency difference pairing units with high authenticity and meeting the mutation feature standard are selected.

[0074] The amplitude change values and frequency difference data are organized into the extracted amplitude change values and frequency difference data, the data of the amplitude change values and frequency difference indices are arranged, the amplitude change values are sorted in ascending order, the corresponding frequency difference index positions are synchronously adjusted using the index positions after sorting the amplitude change values, the frequency difference data is normalized, the frequency difference data is scaled to a set interval, a two-way index table of the normalized results of the amplitude change values and frequency differences is established, and the amplitude change value intervals are batch merged to complete the synchronous organization of the amplitude change values and frequency difference data.

[0075] The specific steps for generating the segment intensity analysis result are as follows:

[0076] Based on the mutation voice segment identification set, extract the frequency values and amplitude values within the segments, read each segment sequentially and extract the frequency change values and amplitude change values of the sampling points within the corresponding segments. Combine with the digital human dynamic expression and posture synchronization algorithm. While extracting the frequency and amplitude changes, synchronize these changes with the actions and expressions of the virtual character to enhance the coordination between the voice and visual effects. Classify the channel numbers to which the sampling points belong, classify the channel numbers of the data content, establish the correspondence between the channel numbers and the data content, and generate a channel classification data list;

[0077] Based on the channel classification data list, calculate the product of the amplitude change value and the frequency change value under each channel number. Match the amplitude change values item by item and multiply them with the corresponding frequency change values. Arrange the segment numbers of the product values in sequence. Combine with the digital human speech synthesis technology. While calculating the product values, optimize the voice waveform. Group and combine them into data sequences according to the channel numbers, integrate the data groups corresponding to the channel numbers, and obtain the product combination sequence;

[0078] Based on the product combination sequence, set and assign weighting coefficients, multiply each product value within each group with the weighting coefficient item by item, accumulate the product of the data and the weighting coefficient, and record the amplitude response value and the frequency response value; Combine with the digital human real-time speech synthesis optimization technology. Optimize the voice output during the weighting process, arrange the response results in the order of the segment numbers, combine the response data corresponding to all segments, and generate the segment intensity analysis result;

[0079] Based on the mutation voice segment identification set, use the data classification mapping algorithm to extract the frequency values and amplitude values within the segments. Use the DataFrame structure of the Pandas library in Python, call the loc function to read each segment sequentially and extract the frequency change values and amplitude change values of the sampling points within the corresponding segments. Classify and process using the channel number field of the sampling points, use the groupby method to classify the data content according to the channel numbers, generate a correspondence table between the channel numbers and the data content, use the reset_index method to rearrange the index, organize and classify the data set, and generate a channel classification data list;

[0080] Based on the channel classification data list, use the vectorized product calculation method to calculate the product of the amplitude change value and the frequency change value under each channel number. Use the multiply function of the NumPy library to perform the multiplication process of matching the amplitude change values with the corresponding frequency change values item by item. Perform a sort_values sorting operation through the number field, arrange the segment numbers corresponding to the product values in sequence, apply the groupby method to aggregate and combine the data according to the channel numbers, use the aggregate(sum) method to integrate the product value set, organize it into a continuous data group, and obtain the product combination sequence;

[0081] Based on the product combination sequence, the weighted cumulative response calculation method is adopted. The weighted coefficient w is set, and the initial weight array is set. The multiply function of the NumPy library is used to perform the item-by-item multiplication of each group's product value and the corresponding weighted coefficient. The sum function is applied for cumulative processing, accumulating the data product term and the weighted coefficient term. The amplitude response value and the frequency response value after accumulation are recorded. All response data are sorted in the order of segment numbers through the sort_index method of DataFrame, and the response data sets corresponding to all segments are combined to generate the segment strength analysis result.

[0082] The specific steps for generating the voice output freeze cache structure are as follows:

[0083] Based on the segment strength analysis result, the channel number and the response value are extracted. The channel number and the corresponding response value are located by the index order positioning method, and the channel number list and the response value list are arranged synchronously. Combining with the digital human dynamic expression and posture synchronization algorithm, when synchronizing the channel number and the response value, the dynamic expression and posture of the virtual character are optimized. The channel number and the response value are matched according to the number mapping method, and a two-way access index relationship is established to generate a two-way mapping table of channel responses;

[0084] Based on the two-way mapping table of channel responses, the corresponding relationship between the channel number and the response strength is established. The index order is rearranged in the ascending order of the channel number. The response value in the reference mapping table is used to assign the number node for the response strength. Combining with the digital human voice synthesis technology, the response strength is optimized and processed, and is organized into a unified structure in the form of a number set. The content of each number and the corresponding response strength are combined to obtain the channel response index set;

[0085] Based on the channel response index set, the freeze state flag is set. The threshold screening method is used to compare the response strength corresponding to the channel number for numerical values, and the channel numbers that meet the freeze conditions are screened out. The freeze flag is set to identify the freeze state. Combining with the digital human real-time voice synthesis optimization technology, the accuracy and stability of the voice output are enhanced in the management of the freeze state. The channel numbers are sequentially written into the storage buffer unit, and the data of each freeze state are combined to generate the voice output freeze cache structure;

[0086] Based on the fragment intensity analysis results, the channel number mapping and index synchronization algorithm is adopted to extract the channel numbers and response values. Using the DataFrame structure of the Pandas library, filtering is performed through the loc function to extract the channel number field and the corresponding response value field. The sort_values method is used to sort in ascending order by channel number, and the reset_index method is used to reset the index order, synchronously arranging the channel number list and the response value list. The mapping relationship between the numbers and values is established through the merge method. The set_index method is used to set the channel number as the primary key, matching the channel number and the response value, and establishing a two-way access index relationship to generate a two-way mapping table of channel responses;

[0087] Based on the two-way mapping table of channel responses, the number index mapping recombination algorithm is adopted to establish the corresponding relationship between the channel numbers and the response intensities. The sort_index method of the DataFrame in the Pandas library is used to rearrange the index order according to the increasing order of the channel numbers. The channel numbers are read one by one through the loc method and assigned values with reference to the corresponding response values. The assign method is used to generate a new response intensity column, and all the channel numbers and response intensity records are sorted in the form of a number set into a unified structure. The concat method is applied to combine the numbers and the corresponding response intensity contents to obtain the channel response index set;

[0088] Based on the channel response index set, the response intensity threshold screening and freeze flag assignment algorithm is adopted. The freeze threshold parameter threshold_value is set to 0.2. The apply method of the DataFrame is used in combination with the lambda expression to compare the response intensity corresponding to the channel number with the threshold, screening out the channel numbers greater than threshold_value. The boolean index is used to mark the freeze state, and the new field freeze_flag is used to assign 1 to indicate freezing and 0 to indicate non-freezing. The channel numbers are sorted in order through sort_values, and the to_parquet function is applied to write to the storage buffer unit in sequence, combining the freeze state data to generate the voice output freeze cache structure.

[0089] The specific method of the channel number and the corresponding response value in the index order positioning method is as follows: for the extracted channel numbers and response values, index the original arrangement order of the channel numbers, map the channel numbers and the corresponding response values, retrieve the response values according to the channel number index order, locate the actual response value position corresponding to the channel number, and establish the synchronous relationship between the channel number index pointer and the response value index pointer to complete the mapping binding from the channel number index to the response value index.

[0090] The specific steps for generating the voice response feature output matrix are as follows:

[0091] Based on the voice output freeze cache structure, obtain the output vector value, sequentially traverse to locate the output value corresponding to the channel number, and retrieve the historical record value corresponding to the channel number. Combining the digital human voice synthesis optimization technology, when obtaining the output vector, optimize the capture accuracy of the voice data, subtract the current output value from the historical record value item by item, calculate the channel number difference, combine the difference data, and generate an output vector difference set;

[0092] Based on the output vector difference set, screen and process the differences, extract the absolute values of the differences item by item, and then compare the absolute values with the set threshold item by item. Screen out the channel numbers whose absolute values of the differences meet the preset values. Combining the digital human real-time voice synthesis technology, optimize the difference screening process, extract the corresponding index number set, sort the screened channel numbers in ascending order, and obtain the screened index number group;

[0093] Based on the screened index number group, arrange the index numbers, organize the index numbers of each channel, arrange the index numbers row by row according to the specified row-column matrix rule. Combining the digital human dynamic expression and posture synchronization algorithm, during the index number sorting process, synchronize the voice output with the dynamic expressions and postures of the virtual characters, combine continuous index segments to form row vectors, stack the row vectors to form the main body of the matrix, output the overall data structure of the matrix, and generate a voice response feature output matrix;

[0094] Based on the voice output freeze cache structure, adopt the vector difference calculation algorithm to obtain the output vector value. Use the array data structure in the NumPy library to structure the output value. Use the for loop combined with enumerate to index and locate the output value corresponding to the channel number. Call the loc method of Pandas to retrieve the historical record value corresponding to the channel number. Subtract the current output value from the historical record value item by item using the NumPy.subtract function. Set the parameter out to specify the output array to store the differences and calculate the channel number difference. Combine the difference data corresponding to all channel numbers, and splice the difference vectors into a unified array through the hstack function to generate an output vector difference set;

[0095] Based on the output vector difference set, adopt the absolute value threshold screening algorithm to screen and process the differences. Use the abs function in the NumPy library to extract the absolute values of the differences item by item. Call the where function combined with the set threshold parameter threshold_value set to 0.1, compare and extract the index numbers that meet the condition of being greater than threshold_value item by item. Use the DataFrame.loc function of Pandas to extract the corresponding channel number index set, sort the screened channel numbers in ascending order through the sort_values method, and use the reset_index method to rearrange the index order. After sorting, obtain the screened index number group;

[0096] Based on the screening of index number groups, the row-column matrix mapping and number reorganization algorithm are adopted to arrange the index numbers, use NumPy.arange to generate a continuous number array, call the reshape method to set the row-column matrix rules, preset each row to contain five index numbers, combine continuous index segments to form a single row vector, and use the vstack function to superimpose multiple row vectors to form the matrix body. Use Pandas.DataFrame to construct the output matrix structure, rearrange the indexes according to the channel numbers and row-column rules, output the overall matrix data structure, and generate the speech response feature output matrix.

[0097] See also Figure 2 , a voice interaction system for a mobile digital human, the system comprising:

[0098] Sound wave feature extraction module: Based on the sound wave input channel data, read the time series content of the sound wave signal, extract the waveform amplitude value of each sampling point, record the corresponding timestamp data, combine the digital human speech synthesis technology, optimize the sampling data through frequency domain transformation processing, improve the accuracy of speech waveform capture, extract the frequency value of each time point through frequency domain transformation processing, calculate the difference between the frequency values ​​of adjacent time points, generate two sets of data sets, namely the amplitude change sequence and the frequency difference sequence, combine the amplitude change sequence and the frequency difference sequence, and obtain the speech signal fluctuation feature group;

[0099] Mutation feature screening module: Based on the voice signal fluctuation feature group, the amplitude change values ​​in the amplitude change sequence that are within the preset amplitude limit range are screened, and the frequency difference values ​​in the frequency difference sequence that are within the preset frequency limit range are screened by using an adversarial neural network. Combined with the digital human emotion analysis model, the emotional fluctuations in the mutation features are further analyzed, and the screened amplitude change values ​​and frequency differences are combined into pairing units to distinguish and classify the generated results and the real data, and the pseudo-mutation samples with large deviations from the real distribution are eliminated. The valid pairing unit set that meets the mutation feature standard is retained, the valid pairing unit set is processed and numbered, the mutation segment index is screened, the mutation segment index set is integrated, and the mutation speech segment identifier set is generated;

[0100] Segment strength analysis module: Based on the mutation voice segment identification set, extract the frequency value and amplitude value in each segment, classify and organize the frequency value and amplitude value of the sound wave input channel number, combine the digital human voice-visual interaction algorithm, optimize the synchronization of voice and dynamic expression of virtual characters in the analysis process, calculate the product of the amplitude change value and the frequency difference in each channel, perform weighted coefficient addition processing on the product result, record the total amplitude response value and the total frequency response value in each channel, combine the total amplitude response value and the total frequency response value, establish the response strength set corresponding to the segment, and generate the segment strength analysis result;

[0101] Response Index Establishment Module: Based on the fragment intensity analysis results, establish a two-way mapping relationship between the channel number and the response intensity, extract the total amplitude response value and the total frequency response value data corresponding to each channel number, index and organize according to the number order to establish a channel index table, set the freeze status flag, and perform write processing on the storage buffer for the freeze flag value; combine the digital human real-time speech synthesis optimization technology to optimize the processing and storage of response data to ensure the smoothness of voice output and the synchronization of virtual character actions; establish a freeze status buffer table structure and generate a voice output freeze cache structure;

[0102] Feature Output Generation Module: Based on the voice output freeze cache structure, extract the output vector value in the current time segment and the output vector value recorded in the corresponding historical time segment, and calculate the difference between the current value and the historical value; combine the digital human speech synthesis optimization technology to enhance the smoothness of the output difference, screen the data points in the difference set whose absolute value is greater than the threshold standard, extract the index numbers of the qualified data points, arrange them in combination according to the index numbers, assemble them into a standard matrix structure, establish a sound wave channel number matrix index, and generate a voice response feature output matrix.

[0103] The above are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A voice interaction method for a mobile digital human, characterized in that, It includes the following steps: Step 1: Based on the input data of the voice acquisition channel, read the time sequence of the acoustic wave signal. By sequentially extracting the waveform amplitude values of the sampling points and recording the corresponding timestamps, extracting the frequency values, and subtracting the adjacent frequency values, generate a voice signal fluctuation feature group; Step 2: Based on the voice signal fluctuation feature group, screen the amplitude and frequency data within the boundary range, generate a mutation candidate feature group, use an adversarial neural network to expand and generate and discriminate valid pairs, obtain a valid mutation pair group, assign numbers to screen mutation segments, and output a mutation voice segment identification set; Step 3: Based on the mutation voice segment identification set, extract the frequency values and amplitude values, classify and sort the data within each segment according to the channel number, obtain the product of the amplitude change value and the frequency difference, sequentially perform weighted coefficient addition processing and record the amplitude response and frequency response values, and generate a segment intensity analysis result; Step 4: Based on the segment intensity analysis result, obtain the response value corresponding to the channel number through bidirectional mapping, establish the index relationship between the channel number and the response intensity, set the freeze state flag and write it into the storage buffer structure, and generate a voice output freeze buffer structure; Step 5: Based on the voice output freeze buffer structure, obtain the difference between the current value and the historical record value of the output vector, screen the difference and locate the vector index number that meets the threshold standard, arrange and assemble the index combination into a matrix structure, and generate a voice response feature output matrix.

2. The voice interaction method of the mobile digital human according to claim 1, wherein The specific steps for generating the voice signal fluctuation feature group are as follows: Based on the input data of the voice acquisition channel, read the time sequence of the acoustic wave signal, extract the acoustic wave amplitude value of the sampling point, use the direct assignment method of the amplitude value to record the timestamp corresponding to the sampling point, combine the sampling point amplitude and timestamp in chronological order, establish a unified sequence structure, and form a waveform amplitude time sequence data group; Based on the waveform amplitude time sequence data group, calculate the amplitude change rate in adjacent time periods, extract the change rate by dividing the difference by the time difference, screen the continuously changing amplitude values, set a threshold range to discriminate the abnormal change rate section, extract it into abnormal segments, and obtain an amplitude change feature group; Based on the amplitude change feature group, extract the main frequency within the corresponding time period, subtract the adjacent frequency values sequentially by using the differential method of the frequency sampling value, combine the frequency differences to form a continuously changing sequence, screen the complete cycle frequency fluctuation change data, and generate a voice signal fluctuation feature group.

3. The voice interaction method of the mobile digital human according to claim 1, characterized in that, The specific steps for generating the mutation voice segment identification set are as follows: Based on the voice signal fluctuation feature group, extract the amplitude change value and the frequency difference, sort out the amplitude change value and the frequency difference data, locate the amplitude change value and the frequency difference, set the boundary extreme values of the amplitude change value and the frequency difference, compare the amplitude change value and the frequency difference, screen out the data pairs that exceed the boundary, establish a data set that meets the boundary conditions, and generate a mutation candidate feature group; Based on the mutation candidate feature group, use the adversarial neural network generator to generate and expand the combined samples of the amplitude change value and the frequency difference, discriminate the authenticity of the generated samples and the original samples, screen out the non-real mutation samples, retain the amplitude change value and the frequency difference pairing units that meet the mutation feature standard, combine the effective unit data groups, and obtain a valid mutation pair group; Based on the effective mutation pairing group, assign sampling segment numbers, mark data numbers using the corresponding index method of amplitude change amplitude and frequency change amplitude, screen the index numbers where the mutation amplitude and mutation frequency match, combine them to form a continuous numbered paragraph, and then organize the number set to generate a mutation voice segment identification set.

4. The voice interaction method of the mobile digital human according to claim 1, characterized in that, The adversarial neural network, according to the formula: Wherein: is the loss function, G is the generator neural network module that receives the input noise vector to generate mutant samples, D is the discriminator neural network module, E is the expectation operator, x is the original real sample, ~ is the sampling symbol from the distribution, p data (x) is the data distribution of the real sample, log is the logarithm operator, D(x) is the discriminant output value of the discriminator for the input real sample x, z is the noise vector, p z (z) is the sampling distribution of the noise variable z, D(G(z)) is the discriminant output value of the discriminator for the sample G(z) generated by the generator, 1 - D(G(z)) is the probability that the generated sample is judged as a false sample, λ is the weight adjustment coefficient, x′ is the mutant sample generated by the generator, p mut (x′) is the sampling distribution of the generated mutant sample x′, x target is the target mutant standard sample, ||x′ - x target || is the Euclidean distance between the generated mutant sample x′ and the target sample x target and, ||x′ - x target || 2 is the square value of the above Euclidean distance.

5. The voice interaction method of the mobile digital human according to claim 3, characterized in that, The process of organizing the amplitude change value and frequency difference data refers to first extracting the amplitude change value and frequency difference data, arranging the amplitude change value and frequency difference data, sorting the amplitude change values in ascending order, synchronously adjusting the corresponding frequency difference index positions using the index positions after sorting the amplitude change values, normalizing the frequency difference data, scaling the frequency difference data to a set interval, establishing a two-way index table for the normalized results of the amplitude change value and frequency difference, and batch-merging the amplitude change value intervals to complete the synchronous organization of the amplitude change value and frequency difference data.

6. The voice interaction method of the mobile digital human according to claim 1, wherein, The specific steps to generate the fragment intensity analysis result are as follows: Based on the mutation voice segment identification set, extract the frequency values and amplitude values within the segment, read and extract the frequency change values and amplitude change values of the sampling points within the corresponding segment one by one, classify the channel numbers to which the sampling points belong, classify the channel numbers of the data content, establish the corresponding relationship between the channel numbers and the data content, and generate a channel classification data list; Based on the channel classification data list, calculate the product of the amplitude change value and the frequency change value under each channel number, match the amplitude change values item by item and multiply them by the corresponding frequency change values, arrange the segment numbers of the product values in order, group them by channel number into data sequences, integrate the data groups corresponding to the channel numbers, and obtain the product combination sequence; Based on the product combination sequence, set and assign weighting coefficients, multiply the product values within each group by the weighting coefficients item by item, accumulate the product of the data and the weighting coefficients, record the amplitude response value and the frequency response value, arrange the response results in the order of the segment numbers, combine the response data corresponding to all segments, and generate the fragment intensity analysis result.

7. The voice interaction method of the mobile digital human according to claim 1, wherein The specific steps to generate the voice output freeze cache structure are as follows: Based on the fragment intensity analysis result, extract the channel numbers and response values, locate the channel numbers and the corresponding response values in the indexing order positioning method, and synchronously arrange the channel number list and the response value list. Match the channel numbers and the response values according to the number mapping method, establish a two-way access index relationship, and generate a channel response two-way mapping table; Based on the channel response two-way mapping table, establish the corresponding relationship between the channel numbers and the response intensities, rearrange the index order using the ascending sorting method of the channel numbers, assign number nodes to the response intensities with reference to the response values in the mapping table, organize them into a unified structure in the form of a number set, combine each number and the corresponding response intensity content, and obtain the channel response index set; Based on the channel response index set, set the freeze status flag, compare the response intensities corresponding to the channel numbers using the threshold screening method, screen out the channel numbers that meet the freeze conditions, set the freeze flag to identify the freeze status, obtain the channel numbers in order and write them into the storage buffer unit in sequence, combine all the freeze status data, and generate the voice output freeze cache structure.

8. The voice interaction method of the mobile digital human according to claim 7, wherein The channel number and the corresponding response value of the index sequence positioning method are specifically as follows: extract the channel number and the response value, index the original arrangement order of the channel numbers, map the channel numbers and the corresponding response values, retrieve the response value according to the index order of the channel numbers, locate the actual position of the response value corresponding to the channel number, and establish a synchronous relationship between the channel number index pointer and the response value index pointer to complete the mapping binding from the channel number index to the response value index.

9. The voice interaction method of the mobile digital human according to claim 1, wherein The specific steps for generating the voice response feature output matrix are as follows: Based on the voice output frozen cache structure, obtain the output vector value, sequentially traverse and locate the output value corresponding to the channel number, retrieve the historical record value corresponding to the channel number, subtract the current output value from the historical record value item by item, calculate the channel number difference, combine the difference data, and generate a set of output vector difference sets; Based on the set of output vector differences, filter and process the differences, extract the absolute value of the difference item by item, then compare the absolute value with the set threshold item by item, filter out the channel numbers whose absolute value of the difference meets the preset value, extract the corresponding index number set, arrange the filtered channel numbers in ascending order, and obtain the filtered index number group; Based on the filtered index number group, arrange the index numbers, organize the index numbers of each channel, arrange the index numbers row by row according to the specified row-column matrix rule, combine consecutive index segments to form a row vector, stack the row vectors to form the main body of the matrix, output the overall data structure of the matrix, and generate the voice response feature output matrix.

10. A voice interaction system for a mobile digital human, characterized in that, According to the voice interaction method of the mobile digital human according to any one of claims 1-9, the system includes: Sound wave feature extraction module: Based on the sound wave input channel data, read the time series content of the sound wave signal, extract the waveform amplitude value of each sampling point, record the corresponding timestamp data, extract the frequency value of each time point through frequency domain transformation processing, calculate the difference between the frequency values of adjacent time points, generate two sets of data, namely the amplitude change sequence and the frequency difference sequence, and combine the amplitude change sequence and the frequency difference sequence to obtain the voice signal fluctuation feature group; Mutation feature screening module: Based on the voice signal fluctuation feature group, screen the amplitude change values in the amplitude change sequence that are within the preset amplitude limit range, and at the same time use the adversarial neural network to screen the frequency differences in the frequency difference sequence that are within the preset frequency limit range. Combine the screened amplitude change values and frequency differences into a pairing unit, distinguish and classify the generated results and the real data, eliminate the pseudo-mutation samples with a large deviation from the real distribution, retain the set of effective pairing units that meet the mutation feature criteria, process the set of effective pairing units and number them, screen the mutation segment indexes, integrate the mutation segment index set, and generate the mutation voice segment identification set; Fragment intensity analysis module: Based on the set of mutated speech fragment identifiers, extract the frequency values and amplitude values within each fragment, classify and organize the frequency values and amplitude values of the acoustic wave input channel numbers, calculate the product results of the amplitude change values and frequency differences within each channel, perform weighted coefficient addition processing on the product results, record the total amplitude response value and total frequency response value within each channel, combine the total amplitude response value and total frequency response value, establish a response intensity set corresponding to the fragment, and generate a fragment intensity analysis result; Response index establishment module: Based on the fragment intensity analysis result, establish a two-way mapping relationship between the channel number and the response intensity, extract the total amplitude response value and total frequency response value data corresponding to each channel number, index and organize according to the number sequence to establish a channel index table, set a freeze status flag, perform a write operation on the storage buffer for the freeze flag value, establish a freeze status buffer table structure, and generate a voice output freeze cache structure; Feature output generation module: Based on the voice output freeze cache structure, extract the output vector value under the current time fragment and the output vector value recorded under the corresponding historical time fragment, calculate the difference between the current value and the historical value, filter the data points in the difference set whose absolute value is greater than the threshold standard, extract the index numbers of the data points that meet the conditions, arrange them in combination according to the index numbers, assemble them into a standard matrix structure, establish an acoustic wave channel number matrix index, and generate a voice response feature output matrix.

Citation Information

Cited By

  • Digital human quality assessment method and device

    CN120726537A

  • Speech synthesis method and system for human-computer interaction

    CN120877706A

  • A speech synthesis method and system for human-computer interaction

    CN120877706B