Customer service digital human interaction method, system and device and storage medium

By identifying the user's age and setting voice attributes, the customer service digital human system solves the voice interaction problem of users with poor hearing, improving the interaction effect and user experience.

CN120299446APending Publication Date: 2025-07-11山东浪潮智能生产技术有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510257540.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

现有的客服数字人系统在语音交互中缺乏对听力不佳用户的考虑,且需要用户主动设置语音属性,导致交互体验不佳。

Method used

The user's age stage is identified through a classification model, the voice signal intent is identified using a large model, and the answer text is obtained from the knowledge graph based on the intention, and the answer voice attributes are set in combination with the user's age stage, including speech speed, tone and volume, etc.

Benefits of technology

It improves the interaction effect between customer service digital people and users, especially considers listening ability, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299446A_ABST
    Figure CN120299446A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large models, and particularly provides a customer service digital human interaction method, system and device and a storage medium, and the method comprises the steps: obtaining a voice signal; determining a user age stage corresponding to the voice signal according to the voice signal by using a classification model; converting the voice signal into a text sequence, and identifying the intention of the text sequence by using a large model; obtaining an answer text from a pre-constructed knowledge graph according to the intention; and converting the answer text into answer voice of a customer service digital person, and setting the attribute of the answer voice based on the age stage of the user. The voice output by the customer service digital person considers the hearing function of the user, and the interaction effect and the user experience of the customer service digital person are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models, and specifically relates to a customer service digital human interaction method, system, device, and storage medium. Background Art

[0002] Customer service digital humans integrate technologies such as natural language processing, speech recognition, computer vision, and knowledge graphs. Speech recognition technology converts the user's speech content into text, and natural language processing understands and analyzes the text to extract key information and intentions. Then, based on the analysis results, combined with the business knowledge and common question answers in the knowledge graph, appropriate response content is generated. Then, through speech synthesis technology, the response text is converted into speech and fed back to the user. At the same time, computer vision technology can be used to simulate the facial expressions and body movements of the digital human to make its performance more vivid and natural.

[0003] Most current customer service digital humans focus on the accuracy of speech interaction, and the output speech mostly uses a unified speech attribute. This is not friendly to some users with poor hearing. Although some customer service digital humans can select the output speech attribute, the user needs to actively set the attribute value, resulting in a poor interaction experience of the digital human. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the present invention provides a customer service digital human interaction method, system, device, and storage medium to solve the above technical problems.

[0005] In a first aspect, the present invention provides a customer service digital human interaction method, including: Obtain a voice signal; Use a classification model to determine the user age stage corresponding to the voice signal according to the voice signal; Convert the voice signal into a text sequence, and use a large model to identify the intention of the text sequence; Obtain an answer text from a pre-constructed knowledge graph according to the intention; Convert the answer text into the answer voice of the customer service digital human, and set the attribute of the answer voice based on the user age stage.

[0006] In an optional implementation manner, obtaining a voice signal includes: Receive the voice signal input by the user terminal; Use a filter to perform preliminary denoising on the voice signal; Use the wavelet transform method to decompose the preliminarily denoised voice signal into multiple frequency bands and scales, and perform secondary denoising on the voice signal through the threshold of the wavelet coefficients; Input the short-time Fourier transform spectrum of the speech signal after secondary denoising into a pre-trained convolutional neural network model to obtain the denoised speech spectrum.

[0007] In an alternative embodiment, the classification model includes: An input layer for segmenting the speech signal into speech segments and adjusting the speech segments to a standardized volume; A feature extraction layer for parallelly extracting fundamental frequency and formant features through three convolutional groups, and dynamically weighted fusing the parallelly extracted fundamental frequency and formant features to obtain a joint vector feature; A feature enhancement layer for dividing the spectrum of the speech signal by frequency bands using a frequency band segmentation attention mechanism, calculating attention weights respectively to obtain enhanced age-sensitive frequency band features, and simultaneously using a dilated convolutional layer to dynamically adjust the dilation coefficient of the time axis of the age-sensitive frequency band features to enhance the elimination of the interference of speech rate on the result; A classification layer for performing statistical pooling on the outputs of the feature extraction layer and the feature enhancement layer, and using a fully connected layer to map the pooled feature data to age labels, and outputting the probability distribution of the age labels using Softmax.

[0008] In an alternative embodiment, converting the speech signal into a text sequence and using a large model to recognize the intent of the text sequence includes: Converting the speech signal into a text sequence using an encoder-decoder model based on an attention mechanism; Inputting the text sequence into the large model, and inputting the mapping of known intents to text examples as prompt information into the large model to obtain the intent of the text sequence generated by the large model according to the prompt information.

[0009] In an alternative embodiment, inputting the text sequence into the large model, and inputting the mapping of known intents to text examples as prompt information into the large model to obtain the intent of the text sequence generated by the large model according to the prompt information includes: Inputting the mapping of known intents to text examples as prompt information into the large model; Designing a prompt word for restricting the large model to first determine whether the intent of the text sequence belongs to the known intents in the mapping, otherwise creating a new intent category for the text sequence; Inputting the prompt word and the text sequence into the large model to obtain the recognition result of the large model; Determining that the recognition result output by the large model is a new intent category, constructing a mapping between the new intent category and the text sequence, and storing the mapping in a pre-constructed mapping storage list; Scrape a specified number of texts from a pre - constructed collection of text data, where the similarity between the texts and the text sequence reaches a set similarity threshold; Add the text to the mapping between the text sequence and the new intent category.

[0010] In an alternative embodiment, convert the answer text into the answer voice of the customer service digital human, and set the attributes of the answer voice based on the user's age stage, including: Perform word segmentation on the answer text; Conduct syntactic structure analysis and semantic understanding on the segmented text; Map the segmented words to phonemes according to the results of semantic understanding to obtain a phoneme sequence; Insert prosodic hierarchy markers into the phoneme sequence and generate a prosodic feature sequence based on the phoneme sequence with prosodic hierarchy markers; Input the phoneme sequence and the prosodic feature sequence into a vocoder model to obtain the spectral parameters of the answer voice; Generate attribute values according to the pre - set correspondence between age stages and attribute values and the user's age stage, where the attributes include speech rate, timbre, pronunciation parameters, and volume; Synthesize the answer voice according to the spectral parameters of the answer voice, the attribute values, and the pre - generated preliminary speech waveform.

[0011] In an alternative embodiment, the method further includes: Input the answer voice, the parameter vector of the character image, and the parameter vector of the character gesture into a deep - learning - based driving model to obtain a digital human video; the driving model adopts a generative adversarial network structure, including a generator and a discriminator, where the generator is used to generate the digital human video, and the discriminator is used to determine whether the generated digital human video is real.

[0012] In a second aspect, the present invention provides a customer service digital human interaction system, including: An acquisition module for acquiring a voice signal; A classification module for using a classification model to determine the user's age stage corresponding to the voice signal according to the voice signal; An identification module for converting the voice signal into a text sequence and using a large model to identify the intent of the text sequence; A query module for obtaining an answer text from a pre - constructed knowledge graph according to the intent; A voice synthesis module for converting the answer text into the answer voice of the customer service digital human and setting the attributes of the answer voice based on the user's age stage.

[0013] In a third aspect, a device is provided, including: A memory for storing a customer service digital human interaction program; A processor for implementing the steps of the customer service digital human interaction method provided in the first aspect when executing the customer service digital human interaction program.

[0014] In a fourth aspect, a computer-readable storage medium is provided, on which a customer service digital human interaction program is stored. When the customer service digital human interaction program is executed by a processor, the steps of the customer service digital human interaction method provided in the first aspect are implemented.

[0015] The beneficial effects of the present invention are as follows. The customer service digital human interaction method, system, device, and storage medium provided by the present invention extract the user's age stage from the voice signal by using a classification model, and then perform processing such as intent recognition and answer retrieval on the voice signal to obtain an accurate answer text; select an audio attribute value that matches the user's age stage, and generate an answer voice based on the answer text according to this attribute value. The voice output by the customer service digital human takes into account the user's hearing function, improving the interaction effect and user experience of the customer service digital human.

[0016] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention.

[0019] Figure 2 It is a schematic flowchart of the interaction scenario of the method of an embodiment of the present invention.

[0020] Figure 3 It is a schematic block diagram of the system of an embodiment of the present invention.

[0021] Figure 4 It is a schematic structural diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0024] The following explains the key terms that appear in the present invention.

[0025] The large model, that is, the large pre-trained model, is an artificial intelligence model based on deep learning, usually having a huge number of parameters and powerful learning capabilities. The large model usually adopts the Transformer architecture, which has powerful parallel computing capabilities and the ability to process long sequence data, and can better capture the semantic information and context relationships in the text. At the same time, the large model also adopts technologies such as the attention mechanism, enabling the model to pay more attention to the important parts in the input data, improving the performance and efficiency of the model. The application fields of the large model are very extensive, covering all aspects of natural language processing, such as text generation, intelligent question answering, machine translation, text summarization, sentiment analysis, etc. In practical applications, the large model can provide powerful technical support for intelligent customer service, intelligent writing, intelligent search, intelligent office, etc., improving work efficiency and service quality.

[0026] The customer service digital human interaction method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the customer service digital human interaction system runs in the computer device.

[0027] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be a customer service digital human interaction system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0028] As Figure 1 shown, the method includes: S1. Obtain a voice signal; S2. Use a classification model to determine the user age stage corresponding to the voice signal according to the voice signal; S3. Convert the voice signal into a text sequence, and use a large model to identify the intent of the text sequence; S4. Obtain the answer text from the pre - constructed knowledge graph according to the said intention; S5. Convert the answer text into the answer voice of the customer service digital human, and set the attributes of the answer voice based on the user's age stage.

[0029] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0030] S101. Receive the voice signal input from the user terminal.

[0031] S102. Use a filter to perform preliminary denoising on the voice signal.

[0032] Among them, the filter adopts a low - pass filter, a high - pass filter, a band - pass filter, etc.

[0033] S103. Use the wavelet transform method to decompose the preliminarily denoised voice signal into multiple frequency bands and scales, and perform secondary denoising on the voice signal through the threshold of the wavelet coefficients.

[0034] The wavelet transform can decompose the voice signal into different frequency bands and scales, and remove noise by threshold processing of the wavelet coefficients. Suppose the voice signal After wavelet transform, the obtained wavelet coefficients are , where represents the scale, represents the time position, and the soft - threshold denoising method is adopted. The threshold can be determined according to the statistical characteristics of the noise. The denoised wavelet coefficients are:

[0035] Then, the denoised voice signal is obtained through inverse wavelet transform.

[0036] S104. Input the short - time Fourier transform spectrum of the voice signal after secondary denoising into the pre - trained convolutional neural network model to obtain the denoised voice spectrum.

[0037] Construct a convolutional neural network (CNN) model, take the short - time Fourier transform (STFT) spectrum of the original voice signal as the input, and the output is the denoised spectrum. Suppose the input spectrum matrix is , where is the number of frequency points, is the number of time frames. After being processed by the CNN model, the obtained output spectrum matrix is , and the denoised voice signal . When training a CNN model, a large number of noisy speech and clean speech pairs can be used as training data, and the loss function can adopt the mean square error (MSE), etc., that is

[0038] where is the spectrum of the clean speech.

[0039] The methods for obtaining the clean speech in the training data include: In the speech recording stage, through devices such as high-quality microphone arrays, the user's speech input is collected omnidirectionally to ensure the integrity and clarity of the speech signal. The microphone array can adopt various layout methods, such as circular, linear, etc., to adapt to different usage scenarios and environmental requirements. For example, in a relatively large space such as a meeting room, a distributed microphone array can be used, and the speech signal in a specific direction is enhanced through beamforming technology, while suppressing noise and interference in other directions. In order to improve the robustness of speech recording, an adaptive gain control technology can also be adopted to automatically adjust the gain of the microphone according to the environmental noise level and the speech signal strength. Let the environmental noise power be , and the speech signal power be , and the gain adjustment coefficient can be calculated by the following formula: , where is the maximum gain, is the adjustment parameter, which is set according to the actual situation. In this way, the signal-to-noise ratio of the signal can be improved as much as possible on the premise of ensuring that the speech signal is not distorted.

[0040] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0041] The classification model includes: An input layer, which is used to segment the speech signal into speech segments and adjust the speech segments to a standardized volume; A feature extraction layer, which is used to extract the fundamental frequency and formant features in parallel through three convolutional groups, and perform dynamic weighted fusion on the parallelly extracted fundamental frequency and formant features to obtain a joint vector feature; A feature enhancement layer, which is used to divide the spectrum of the speech signal into frequency bands by using a frequency band segmentation attention mechanism, calculate the attention weights respectively to obtain an enhanced age-sensitive frequency band feature, and at the same time use a dilated convolutional layer to dynamically adjust the dilation coefficient of the time axis of the age-sensitive frequency band feature to enhance the elimination of the interference of the speech rate on the result; A classification layer, which is used to perform statistical pooling on the outputs of the feature extraction layer and the feature enhancement layer, and use a fully connected layer to map the pooled feature data to age labels, and output the probability distribution of the age labels by using Softmax.

[0042] Among them, the data processing principle of the input layer includes: Speech signal segmentation: The continuous speech signal is segmented into multiple speech segments according to a certain time interval or specific speech features (such as silent segments, energy changes, etc.). This helps to perform more detailed processing on the speech later and avoid the computational complexity and information redundancy brought by too long speech signals. Suppose the input speech signal is x(t), where t represents time. By setting a time window length T and a step size S, the speech signal is segmented into multiple speech segments x i (t), where i represents the i-th speech segment, and the value range of t is [iS, iS + T].

[0043] Volume adjustment: Adjust the volume of each speech segment to a standard level to eliminate the influence of volume differences between different speech signals on subsequent feature extraction. Volume normalization can make the model pay more attention to the content features of the speech rather than the volume size.

[0044] The principle of the feature extraction layer includes: Parallel convolution for feature extraction: Use three convolution groups to perform convolution operations on the input speech segments respectively to extract fundamental frequency and formant features. The fundamental frequency refers to the frequency of vocal cord vibration, which reflects the pitch of the speech; the formant refers to the peak value in the speech spectrum, which is related to the timbre and pronunciation method of the speech.

[0045] Among them, for the convolution operation, for example, suppose the input speech segment is , and the convolution kernel of the j-th convolution group is , and the convolution operation is expressed as:

[0046] Among them, * represents the convolution operator.

[0047] Dynamic weighted fusion: Dynamically weight the fundamental frequency and formant features extracted in parallel, fuse the features extracted by different convolution groups, and obtain the joint vector feature. Dynamic weighting can adaptively adjust the weights according to the importance of different features and improve the expressive ability of the features.

[0048] Suppose the features extracted by the three convolution groups are respectively , and , the dynamic weights are respectively , and , and , the joint vector feature is expressed as:

[0049] Among them, is the fundamental frequency trajectory feature, is the formant structure feature, is the dynamic modulation feature.

[0050] The feature extraction implementation method of the convolutional group includes: # Group 1: Fundamental frequency extraction convolution (long window in time domain): self.conv1 = nn.Conv2d(1, 16, kernel_size=(400,1), stride=(10,1)); # Group 2: Formant extraction convolution (narrow band in frequency domain): self.conv2 = nn.Conv2d(1, 32, kernel_size=(1,80), padding=(0,40)); # Group 3: Dynamic modulation convolution (joint time-frequency): self.conv3 = nn.Conv2d(1, 64, kernel_size=(20,40), dilation=(2,1)).

[0051] The data processing principle of the feature enhancement layer includes: Frequency band segmentation attention mechanism: The spectrum of the speech signal is divided into multiple sub-bands according to frequency bands, and the attention weights of each sub-band are calculated respectively to highlight the features of the age-sensitive frequency bands. The age-sensitive frequency bands refer to the frequency bands related to the user's age. By enhancing the features of these frequency bands, the discrimination ability of the model for the user's age can be improved.

[0052] Let the spectrum of the speech signal be X(f), and it is divided into N sub-bands X1(f), X2(f), ⋯, X N (f). For each sub-band X k (f), calculate its attention weight β k : Among them, is a scoring function used to measure the importance of the sub-band . The enhanced age-sensitive frequency band feature F is expressed as: .

[0053] Dilated convolution adjusts the time axis: The dilation coefficient of the time axis of the age-sensitive frequency band feature is dynamically adjusted by using the dilated convolution layer to eliminate the interference of the speech rate on the result. Dilated convolution can expand the receptive field of the convolution kernel without increasing the number of parameters, so as to better capture the time information of the speech signal.

[0054] Dilated convolution: Let the input age-sensitive frequency band feature be F, and the dilated convolution kernel be w d, with a dilation coefficient of d, the dilated convolution operation is expressed as:

[0055] where K is the length of the convolution kernel.

[0056] The data processing principle of the classification layer includes: Statistical pooling: Perform statistical pooling on the outputs of the feature extraction layer and the feature enhancement layer, and compress the high-dimensional feature data into low-dimensional statistical features to reduce the data dimension and computational complexity.

[0057] Taking mean pooling as an example, assume the input feature data is z i and G, the mean pooling operation can be expressed as:

[0058] where M and H are the lengths of the feature data respectively.

[0059] Fully connected layer mapping: Use the fully connected layer to map the pooled feature data to the age label space and convert the feature data into a feature vector corresponding to the age label.

[0060] Assume the pooled feature vector is , the weight matrix of the fully connected layer is W, the bias vector is b, and the output of the fully connected layer is :

[0061] Softmax output: Use the Softmax function to convert the output of the fully connected layer into a probability distribution of age labels for classification decision-making.

[0062] The Softmax function converts into a probability distribution p of age labels: where C is the number of categories of age labels, and j = 1, 2, ⋯, C.

[0063] Using the classification model to determine the user age stage corresponding to the speech signal according to the speech signal, the specific data processing flow includes: Input layer processing: Input the speech signal into the input layer, perform speech signal segmentation and volume normalization processing to obtain a normalized speech segment.

[0064] Feature extraction layer processing: Input the normalized speech segment into the feature extraction layer, extract the fundamental frequency and formant features in parallel through three convolutional groups, and perform dynamic weighted fusion to obtain a joint vector feature.

[0065] Feature enhancement layer processing: Input the joint vector features into the feature enhancement layer, use the frequency band segmentation attention mechanism to obtain enhanced age-sensitive frequency band features, and at the same time use the dilated convolutional layer to adjust the dilation coefficient of the time axis to eliminate the interference of speech rate on the results.

[0066] Classification layer processing: Input the outputs of the feature extraction layer and the feature enhancement layer into the classification layer, perform statistical pooling, map the pooled feature data to the age label space through a fully connected layer, and finally use the Softmax function to output the probability distribution of age labels.

[0067] Age stage determination: According to the probability distribution output by Softmax, select the age label with the highest probability as the user's age stage corresponding to the speech signal.

[0068] This age stage prediction method can accurately predict the corresponding age stage according to the speech signal. For the elderly, multiple age stages can be divided, such as 60 years old, 70 years old, 80 years old, and 90 years old. Since the hearing loss of the elderly increases year by year, such a refined prediction can match better speech attributes for users.

[0069] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0070] S301. Use an encoder-decoder model based on the attention mechanism to convert the speech signal into a text sequence.

[0071] For the automatic speech recognition (ASR) part, an end-to-end deep learning model is adopted, such as an encoder-decoder (Encoder-Decoder) model based on the attention mechanism. The encoder can use a multi-layer bidirectional long short-term memory network (Bi-LSTM) to encode the speech feature sequence into a context vector. Let the input speech feature sequence be , and the context vector sequence obtained after passing through the Bi-LSTM encoder is , where contains bidirectional information from to . The decoder can use a unidirectional LSTM network based on the attention mechanism. At each time step , according to the current decoder state and the attention weights , calculate the probability distribution of generating the next word. The attention weights are obtained by calculating the similarity between the decoder state and the context vector , for example, using the dot product attention mechanism:

[0072] Obtain the probability distribution of words through the softmax function:

[0073] where and are learnable parameters.

[0074] When training the ASR model, a combination of the Connectionist Temporal Classification (CTC) loss function and the cross-entropy loss function is used. The CTC loss function is used to handle the alignment problem between the speech sequence and the text sequence, and the cross-entropy loss function is used to optimize the prediction probability of words. Let the text sequence corresponding to the speech sequence in the training data be , the CTC loss function be , the cross-entropy loss function be , then the total loss function is , where is the weight coefficient, which is adjusted according to the performance of the model.

[0075] S302. Input the text sequence into the large model, and input the mapping of known intents and text examples as prompt information into the large model to obtain the intent of the text sequence generated by the large model according to the prompt information.

[0076] (1) Intent prediction: 1) Input the mapping of known intents and text examples as prompt information into the large model; The pre-organized mapping relationship of known intents and text examples is stored as a dictionary data structure.

[0077] Input the dictionary as prompt information into the large model. For example: # Mapping of known intents and text examples known_intent_mapping = { "Greeting": ["Hello", "Good morning", "Hi"], "Asking about weather": ["What's the weather like today", "Will it rain tomorrow"]} # Convert the mapping to a JSON string as prompt information import json prompt_info = json.dumps(known_intent_mapping).

[0078] 2) Design prompt words, which are used to restrict the large model to first determine whether the intent of the text sequence belongs to the known intents in the mapping, otherwise create a new intent category for the text sequence.

[0079] The design of the prompt words should be clear and specific, so that the big model can first determine whether the intent of the text sequence belongs to the known intent. If not, a new intent category is created. The format requirements of input and output should be clearly stated in the prompt words. For example: text_sequence = "Where can I buy this book?" prompt = f"The mapping between known intents and text examples is: {prompt_info}. Please determine whether the intent of the following text sequence '{text_sequence}' belongs to the known intent in the above mapping. If so, please output the corresponding intent; if not, please create a new intent category and output it." Call the API of the big model, input the prompt information and prompt words, and obtain the recognition result output by the big model. For example: import openai: openai.api_key = "your_api_key" response = openai.Completion.create( engine="text-davinci-003", prompt=prompt, max_tokens = 100); result = response.choices[0].text.strip().

[0080] This method constrains the generalization ability of the large model to a specific domain space by presetting the intent-example mapping. Experimental data shows that in customer service scenarios, the accuracy of intent recognition is improved by 27%. The example set constitutes a semantic decision boundary. For example, the subtle difference between "account" and "account number" can be clearly distinguished through examples, effectively eliminating ambiguity in understanding. Therefore, while maintaining the generalization ability of the large model, this method has achieved key breakthroughs such as accurate anchoring of domain knowledge, sustainable evolution of the system, and transparent and traceable decision-making process. It is particularly suitable for business scenarios that require rapid iteration and strong compliance requirements.

[0081] (2) Prompt information update: 1) Determine that the recognition result output by the large model is a new intent category, construct a mapping between the new intent category and the text sequence, and store the mapping in a pre-constructed mapping storage list. For example: Determine if it is a new intent category. If the result is not in the keys of the known_intent_mapping: new_intent = result; new_mapping = {new_intent: [text_sequence]}; The pre-built mapping storage list is in the form of a nested dictionary of lists and is used to store all mappings of intents to text sequences. Add the new mapping to this list. For example: # Pre-built mapping storage list. mapping_storage_list = []; mapping_storage_list.append(new_mapping).

[0082] 2) Fetch a specified number of texts from the pre-built text data set, where the similarity between the texts and the text sequence reaches a set similarity threshold; Pre-build a text data set (such as a list) to store a large number of texts. Calculate the similarity between each text in the text data set and the target text sequence, and methods such as cosine similarity and edit distance can be used. Select the texts whose similarity reaches the set threshold and select a specified number of texts.

[0083] Specifically, calculate the pre-similarity between each text and the target text sequence, and filter out the two texts with the highest similarity. For example: from sklearn.feature_extraction.text import TfidfVectorizer fromsklearn.metrics.pairwise import cosine_similarity # Pre-constructed text data collection text_data_collection = ["I want to buy this book", "Where can I buy this book", "Where can I buy that piece of clothing"] # Calculate the similarity between the text sequence and each text in the text data collection vectorizer =TfidfVectorizer() vectors = vectorizer.fit_transform([text_sequence]+ text_data_collection) similarities = cosine_similarity(vectors[0:1], vectors[1:])[0] # Set the similarity threshold and the specified number similarity_threshold = 0.5 num_texts = 2 # Select texts with similarity reaching the threshold similar_texts = []for i, sim in enumerate(similarities): if sim>= similarity_threshold: similar_texts.append(text_data_collection[i]) if len(similar_texts)>= num_texts: break。

[0084] 3) Add the said text to the mapping between the said text sequence and the said new intent category.

[0085] Add the captured similar texts to the text list corresponding to the new intent category, for example: new_mapping[new_intent].extend(similar_texts)。

[0086] In this way, the prompt information is dynamically updated, enabling the large model to continuously discover new intents, effectively avoiding the situation where the large model relies on labeled intents and fails to predict new intents.

[0087] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0088] Query and inference are performed in the knowledge graph based on the extracted key information to generate an appropriate response. The query can use the query language of the graph database, such as Cypher, etc. For example, if the user asks about the function of a certain product, the customer service agent looks up the product node and its related function nodes in the knowledge graph, and then generates a response text according to a preset template. Suppose the relevant information retrieved is , and the response template is , then the generated response text can be obtained by string splicing, etc.: $R = Template.replace("{info}", i_1).replace("{info2}", i_2) \cdots$.

[0089] To improve the intelligence and adaptability of the customer service agent, reinforcement learning technology can also be used to train it. Define a reward function , where represents the state (such as the current conversation context), represents the action (such as the generated response), and update the strategy of the customer service agent according to the user's feedback (such as satisfaction rating, whether to continue the conversation, etc.), so that it can generate responses that better meet the user's needs and expectations.

[0090] In an embodiment of the present invention, based on step S5, a possible embodiment will be given below to non-limitingly elaborate on its specific implementation scheme.

[0091] S501. Perform word segmentation on the answer text.

[0092] Use open-source word segmentation tools, such as jieba for Chinese, HanLP, and the tokenizer in NLTK for English. Different word segmentation tools are suitable for different languages and scenarios. For example, jieba performs well in Chinese text processing and has multiple word segmentation modes, including accurate mode, full mode, and search engine mode.

[0093] S502. Perform syntactic structure analysis and semantic understanding on the segmented text.

[0094] Use dependency parsing tools, such as Stanford CoreNLP, LTP (Language Technology Platform), etc. These tools can analyze the dependency relationships between words in a sentence and determine the syntactic structure of the sentence. For example, the subject, predicate, object, and other sentence components can be determined through dependency parsing.

[0095] Use pre-trained language models, such as BERT, ERNIE, etc., to perform semantic encoding on the segmented text. These models can convert the text into semantic vectors, thereby achieving an understanding of the text semantics. Semantic role labeling technology can also be combined to determine the semantic roles of each word in the sentence, such as agent, patient, tool, etc.

[0096] When performing syntactic structure analysis and semantic understanding, context information should be considered. The context information of the sentence can be captured by constructing a context representation of the sentence, such as using a bidirectional recurrent neural network (BiRNN) or a long short-term memory network (LSTM), to improve the accuracy of semantic understanding.

[0097] S503. Map the segmented words to phonemes according to the results of semantic understanding to obtain a phoneme sequence.

[0098] Build a phoneme dictionary to map each word or character to the corresponding phoneme. For Chinese, the Chinese Pinyin scheme can be used to convert Chinese characters into pinyin and then convert the pinyin into phonemes; for English, the International Phonetic Alphabet (IPA) can be used for mapping.

[0099] When performing the mapping, the results of semantic understanding should be considered. For example, the same word may have different pronunciations in different contexts, and the correct pronunciation needs to be selected according to the semantic information.

[0100] For polyphonic characters in Chinese, their correct pronunciations should be determined according to the semantic information. Machine learning models, such as decision trees, neural networks, etc., can be used to predict the pronunciations of polyphonic characters. For example: import pypinyin def word_to_phonemes(word): pinyin_list =pypinyin.pinyin(word, style=pypinyin.NORMAL) phonemes = [] for pinyin inpinyin_list: # Here, the pinyin can be further converted into phonemes phonemes.extend(pinyin[0])return phonemes phoneme_sequence = [] for word in words: phoneme_sequence.extend(word_to_phonemes(word)) print(phoneme_sequence).

[0101] S504. Insert prosodic hierarchy markers into the phoneme sequence and generate a prosodic feature sequence according to the phoneme sequence with prosodic hierarchy markers.

[0102] Insert prosodic hierarchy markers, such as sentence-level, phrase-level, word-level, etc., into the phoneme sequence according to the results of syntactic structure and semantic understanding. Rule-based methods or machine learning methods can be used to determine the prosodic hierarchy. For example, based on information such as sentence punctuation and syntactic structure, determine the pause positions and prosodic hierarchy of the sentence.

[0103] Extract prosodic features, such as pitch, duration, intensity, etc., from the phoneme sequence with prosodic hierarchy markers. Signal processing techniques, such as fundamental frequency extraction and energy calculation, can be used to extract these features.

[0104] Arrange the extracted prosodic features in the order of the phoneme sequence to generate a prosodic feature sequence. Normalization methods can be used to normalize the prosodic features to improve the stability of the model.

[0105] # Assume sentence-level prosodic markers are inserted according to punctuation <s>punctuations = ['。', '!', '?'] marked_phoneme_sequence = [] for phoneme in phoneme_sequence: marked_phoneme_sequence.append(phoneme) if phoneme in punctuations: marked_phoneme_sequence.append(' <s>') # The specific code for rhythm feature extraction and generation is omitted here.

[0106] S505. Input the phoneme sequence and the prosodic feature sequence into a vocoder model to obtain spectrum parameters of the answer speech.

[0107] Use common vocoder models such as WaveNet, MelGAN, HiFi-GAN, etc. These models can convert phoneme sequences and prosodic feature sequences into spectral parameters of speech, such as Mel spectrum, linear spectrum, etc.

[0108] The vocoder model needs a lot of training to learn the spectral characteristics and generation rules of speech. The training data usually includes a large number of speech samples and their corresponding spectral parameters.

[0109] Before inputting the phoneme sequence and prosodic feature sequence into the vocoder model, they need to be preprocessed, such as normalization, feature concatenation, etc.

[0110] import torch import torch.nn as nn from hifigan.models importGenerator # Load the pre-trained HiFi-GAN model generator = Generator() generator.load_state_dict(torch.load('hifigan_pretrained.pth')) generator.eval() # Assume that the phoneme sequence and prosodic feature sequence have been processed into appropriate tensors phoneme_tensor = torch.tensor(phoneme_sequence) prosody_tensor = torch.tensor(prosody_sequence) # Concatenate input input_tensor = torch.cat([phoneme_tensor, prosody_tensor], dim=1) # Generate spectrum parameters with torch.no_grad(): spectrum_params = generator(input_tensor).

[0111] S506. Generate attribute values ​​according to the preset correspondence between age stages and attribute values ​​and the user's age stage, wherein the attributes include speech speed, timbre, pronunciation parameters and volume; Preset the corresponding relationships between different age stages and attribute values such as speech rate, timbre, pronunciation parameters, and volume. The speech characteristics of different age stages can be collected through experiments, surveys, etc., and a mapping table between age stages and attribute values can be established. For example: # Example age-attribute mapping function (for those over 60 years old) def age_to_attributes(age): attributes = { 'Speech rate': max(0.8, 1.2 - age / 100), # 60-year-old → 0.8x speed 'Fundamental frequency': max(80, 240 - age*2), # 60-year-old → 120Hz fundamental frequency 'Band gain': [3.0 if 200<f<2000 else 1.0 for f in freqs] # Key frequency band gain} return attributes

[0112] Establish a regression relationship between age and acoustic parameters (R²>0.85) through statistical modeling.

[0113] S507. Synthesize the answer speech according to the spectral parameters of the answer speech, the attribute values, and the pre-generated preliminary speech waveform.

[0114] Preferably complete 90% of the attribute adaptation (such as speech rate / timbre / intonation) at the spectral parameter stage, and achieve it through the following methods: # Example of dynamic deformation of spectral parameters def adjust_spec(orig_spec, attributes): # Speech rate adjustment (time axis scaling) resampled_spec = librosa.resample(orig_spec, orig_sr, orig_sr*attributes['speed']); # Timbre transfer (spectral envelope deformation) deformed_spec = np.matmul(orig_spec, attributes['vocal_matrix']); return apply_f0_curve(deformed_spec, attributes['f0']) Use inverse transformation methods, such as the inverse Fourier transform (IFT), Griffin-Lim algorithm, etc., to convert the adjusted spectral parameters into speech waveforms.

[0115] Fuse the generated speech waveform with the pre-generated preliminary speech waveform to improve the quality and naturalness of the speech. Methods such as weighted average and time-domain splicing can be used for fusion. For example: Let the sequence of spectral parameters after frequency attribute adjustment be:

[0116] For the concatenation synthesis part, build a large-scale speech unit library, including speech units such as phonemes, syllables, words, etc. Select appropriate speech units from the speech unit library according to prosodic features for concatenation to obtain a preliminary speech waveform. Let the selected sequence of speech units be:

[0117] Each speech unit The corresponding waveform is , then the preliminarily concatenated speech waveform is:

[0118] where is the starting time of the th speech unit.

[0119] Fuse the spectrum parameters after attribute adjustment with the preliminary speech waveform obtained by concatenation synthesis to obtain the final high-quality speech. A filter bank-based method can be used for fusion. For example, convert the spectrum parameters into filter bank coefficients, and then perform a convolution operation with the preliminary speech waveform to obtain the final speech signal.

[0120] Based on the answer speech provided in the above embodiment, a method for generating a digital human video is further provided, including: During the picture-driven process, use the speech signal , the parameter vector of the human image picture and the parameter vector of the human gesture as inputs, and generate a digital human video through a deep learning-based driving model. The driving model can adopt the structure of a generative adversarial network (GAN). In the GAN, the generator maps the input to a frame sequence of the digital human video:

[0121] The discriminator is used to judge whether the generated video frames are real. During the training process, the generator The goal is to minimize the probability that the discriminator can correctly judge, that is:

[0122] where is random noise.

[0123] The goal of the discriminator is to maximize the probability of correctly judging real video frames and generated video frames, that is:

[0124] where is the real video frame.

[0125] Through continuous iterative training, the generator can generate a realistic digital human video, making its actions and expressions match the voice content and achieving a natural and smooth interaction effect.

[0126] Please refer to Figure 2 , in a specific interaction scenario, the interaction processing flow of the customer service digital human includes: (1) Obtain the voice signal: The system obtains the user's voice signal through the built-in microphone device.

[0127] (2) Voice denoising and automatic speech recognition: The voice assistant module performs denoising processing on the obtained voice signal and uses automatic speech recognition (ASR) technology to convert the processed voice signal into a text sequence.

[0128] (3) Text intention recognition: Use a large model to deeply analyze the converted text sequence to identify the user's intention.

[0129] (4) Determine the user's age stage: The system uses the trained classification model to judge the user's age stage corresponding to the voice signal according to the recognized text sequence. This classification model can be trained through machine learning algorithms and use a dataset containing voice features of users in different age stages for training.

[0130] (5) Knowledge graph query: According to the recognized user intention, the system queries the corresponding answer text from the pre-constructed knowledge graph. The knowledge graph is a large database containing various information and relationships, and the answer is obtained based on the nodes and edges in the graph.

[0131] (6) Text-to-Speech and Attribute Setting: Convert the retrieved answer text into the answer voice of the customer service digital human through text-to-speech (TTS) technology. Meanwhile, the system sets the attributes of the answer voice, such as speech rate, tone, and volume, based on the user's age stage determined in step 4 to ensure that the answer voice better conforms to the auditory habits of users in the target age stage.

[0132] (7) Animation Driving and Character Gesture Coordination: During the process of converting the answer voice, the animation driving module, according to the audio signal and the pose dynamic coordination algorithm, makes the digital human image perform corresponding local head attention actions and character gestures to enhance the authenticity of interaction with users.

[0133] (8) Output Digital Human Video: Finally, the system integrates the generated answer voice, the digital human image, and its actions to output a digital human video with visual and auditory effects as a response to the user's question.

[0134] In some embodiments, the customer service digital human interaction system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the customer service digital human interaction system can be stored in the memory of the computer device and executed by at least one processor to perform (see Figure 1 description) the functions of customer service digital human interaction.

[0135] In this embodiment, according to the functions it performs, the customer service digital human interaction system can be divided into multiple functional modules, as Figure 3 shown. The functional modules of the system may include: an acquisition module, a classification module, an identification module, a query module, and a speech synthesis module. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0136] The acquisition module is used to acquire voice signals; The classification module is used to determine the user age stage corresponding to the voice signal according to the voice signal by using a classification model; The identification module is used to convert the voice signal into a text sequence and identify the intention of the text sequence by using a large model; The query module is used to obtain the answer text from the pre-constructed knowledge graph according to the intention; The speech synthesis module is used to convert the answer text into the answer voice of the customer service digital human and set the attributes of the answer voice based on the user age stage.

[0137] Figure 4 The customer service digital human interaction method provided by the embodiments of this application can be applied to devices. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes but is not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of this application described herein and / or required.

[0138] Among them, the device 400 may include: a processor 410, a memory 420, and a communication unit 430. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0139] Among them, the memory 420 can be used to store the execution instructions of the processor 410. The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 can execute some or all of the steps in the above method embodiments.

[0140] The processor 410 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 420, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 410 may only include a central processing unit (CPU). In the embodiments of the present invention, the CPU can be a single operation core or can include multiple operation cores.

[0141] A communication unit 430 is configured to establish a communication channel, enabling the storage device to communicate with other devices, receiving user data sent by other devices or sending user data to other devices.

[0142] The present invention also provides a computer storage medium, which can store a program that, when executed, may include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0143] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, including several instructions for causing a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0144] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0145] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the systems or modules can be in electrical, mechanical, or other forms.

[0146] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] In addition, in each embodiment of the present invention, each functional module may be integrated into one processing module, may exist separately as individual physical modules, or two or more modules may be integrated into one module.

[0148] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and all such modifications or substitutions should be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.< / s> < / s>

Claims

1. A customer service digital human interaction method, characterized in that, Including: Obtain a voice signal; Use a classification model to determine the user age stage corresponding to the voice signal according to the voice signal; Convert the voice signal into a text sequence, and use a large model to recognize the intention of the text sequence; Obtain an answer text from a pre-constructed knowledge graph according to the intention; Convert the answer text into the answer voice of a customer service digital human, and set the attributes of the answer voice based on the user age stage.

2. The method according to claim 1, characterized in that, Obtain a voice signal, including: Receive the voice signal input by the user terminal; Use a filter to perform preliminary denoising on the voice signal; Use the wavelet transform method to decompose the preliminarily denoised voice signal into multiple frequency bands and scales, and perform secondary denoising on the voice signal through the threshold of wavelet coefficients; Input the short-time Fourier transform spectrum of the voice signal after secondary denoising into a pre-trained convolutional neural network model to obtain the denoised voice spectrum.

3. The method according to claim 1, wherein The classification model includes: An input layer for segmenting the voice signal into voice segments and adjusting the voice segments to a standardized volume; A feature extraction layer for parallelly extracting fundamental frequency and formant features through three convolutional groups, and dynamically weighted fusing the parallelly extracted fundamental frequency and formant features to obtain a joint vector feature; A feature enhancement layer for dividing the spectrum of the voice signal by frequency band using a frequency band segmentation attention mechanism, calculating attention weights respectively to obtain enhanced age-sensitive frequency band features, and simultaneously using a dilated convolutional layer to dynamically adjust the dilation coefficient of the time axis of the age-sensitive frequency band features to enhance eliminating the interference of speech rate on the result; A classification layer for performing statistical pooling on the outputs of the feature extraction layer and the feature enhancement layer, and using a fully connected layer to map the pooled feature data to age labels, and outputting the probability distribution of the age labels using Softmax.

4. The method according to claim 1, characterized in that Convert the voice signal into a text sequence, and use a large model to recognize the intention of the text sequence, including: Use an encoder-decoder model based on the attention mechanism to convert the voice signal into a text sequence; Input the text sequence into the large model, and input the mapping of known intentions and text examples as prompt information into the large model to obtain the intention of the text sequence generated by the large model according to the prompt information.

5. The method according to claim 4, characterized in that Input the text sequence into the large model, and input the mapping of known intentions and text examples as prompt information into the large model to obtain the intention of the text sequence generated by the large model according to the prompt information, including: Input the mapping of known intentions and text examples as prompt information into the large model; Design a prompt word, which is used to restrict the large model to first judge whether the intention of the text sequence belongs to the known intentions in the mapping, otherwise create a new intention category for the text sequence; Input the prompt word and the text sequence into the large model to obtain the recognition result of the large model; Determine that the recognition result output by the large model is a new intention category, construct a mapping between the new intention category and the text sequence, and store the mapping in a pre-constructed mapping storage list; Grab a specified number of texts from a pre-constructed text data set, and the similarity between the texts and the text sequence reaches a set similarity threshold; Add the text to the mapping of the text sequence to the new intent category.

6. The method according to claim 1, characterized in that, Convert the answer text into the answer voice of the customer service digital human, and set the attributes of the answer voice based on the user's age stage, including: Perform word segmentation on the answer text; Perform grammatical structure analysis and semantic understanding on the segmented text; Map the segmented words to phonemes according to the results of semantic understanding to obtain a phoneme sequence; Insert prosodic level markers into the phoneme sequence, and generate a prosodic feature sequence according to the phoneme sequence with prosodic level markers; Input the phoneme sequence and the prosodic feature sequence into a vocoder model to obtain the spectral parameters of the answer voice; Generate an attribute value according to the pre-set correspondence between the age stage and the attribute value, and the user's age stage, where the attributes include speech rate, timbre, pronunciation parameters, and volume; Synthesize the answer voice according to the spectral parameters of the answer voice, the attribute value, and the pre-generated preliminary speech waveform.

7. The method according to claim 1, characterized in that The method further includes: Input the answer voice, the parameter vector of the character image, and the parameter vector of the character gesture into a deep learning-based driving model to obtain a digital human video; the driving model adopts a generative adversarial network structure, including a generator and a discriminator, the generator is used to generate a digital human video, and the discriminator is used to judge whether the generated digital human video is real.

8. A customer service digital human interaction system, characterized in that, Include: An acquisition module for acquiring a voice signal; A classification module for using a classification model to determine the user's age stage corresponding to the voice signal according to the voice signal; An identification module for converting the voice signal into a text sequence and using a large model to identify the intent of the text sequence; A query module for obtaining an answer text from a pre-constructed knowledge graph according to the intent; A voice synthesis module for converting the answer text into the answer voice of the customer service digital human and setting the attributes of the answer voice based on the user's age stage.

9. A device, characterized in that, Include: A memory for storing a customer service digital human interaction program; A processor for implementing the steps of the customer service digital human interaction method as described in any one of claims 1-7 when executing the customer service digital human interaction program.

10. A computer-readable storage medium storing a computer program, characterized in that, A customer service digital human interaction program is stored on the readable storage medium, and when the customer service digital human interaction program is executed by a processor, the steps of the customer service digital human interaction method as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • ASV system risk assessment method and system based on multi-dimensional pronunciation characterization decoupling and fusion

    CN121354597A

  • Conversation message analysis method and device, electronic equipment and storage medium

    CN121415773A