Intelligent voice interaction method and system based on semantic recognition and call linkage

By extracting temporal features and performing acoustic decoding on the real-time audio stream of multi-party voice calls, and combining this with conditional random fields for semantic recognition, prompt audio carrying harmonic distribution is generated. This solves the problem that voice assistants cannot capture implicit intentions in multi-party calls, and achieves seamless intelligent voice interaction.

CN122637784APending Publication Date: 2026-08-25SHENZHEN CHENGXUNDE COMM SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610887056.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing voice assistants struggle to analyze user conversation semantics in real time, capture users' implicit intentions, and provide voice prompts through accompanying response mechanisms in multi-party conversation contexts, resulting in fragmented voice interaction and disrupted call continuity.

Method used

By acquiring real-time audio streams of multi-party voice calls, temporal features are extracted and acoustic decoding is performed. Combined with context window annotation of sentence boundaries and semantic slots of historical rounds, fuzzy word frequency and intent category entity recognition and dependency extraction are performed using conditional random fields to generate prompt audio carrying harmonic distribution. Accompanying response is achieved through frequency band shifting and mixing superposition.

Benefits of technology

It improves the accuracy of extracting user intent and key information, reduces the probability of false responses due to misidentification, and achieves seamless integration of voice prompts and call audio, enabling imperceptible intelligent interaction in natural scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637784A_ABST
    Figure CN122637784A_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent voice interaction method and system based on semantic recognition and call linkage, belongs to the technical field of voice interaction.The method first acquires the real-time audio stream of multi-party voice call and carries out timing feature extraction and acoustic decoding, combines context window marking sentence boundary and historical semantic slot, generates text sequence;Then, the conditional random field is used to identify ambiguous word frequency and intent category, extract entity and dependency relationship, obtain semantic feature vector carrying semantic slot and trigger condition;Subsequently, the target information requirement is determined by comprehensive analysis of historical confidence and trigger condition;Then, the matching field value and interpretation item are retrieved from the structured knowledge base to generate response text;Through voice synthesis, prompt audio with harmonic distribution is constructed, and frequency band translation is carried out according to the spectral coincidence rate of the prompt audio and the real-time audio stream, and the accompanying response audio is generated by mixing and superimposing and output, the application realizes low interference, high reliable intelligent voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of voice interaction technology, specifically an intelligent voice interaction method and system based on semantic recognition and call linkage. Background Technology

[0002] Intelligent voice interaction technology occupies a crucial position in modern communication networks, serving as a key hub for improving user communication efficiency and information access convenience. However, current voice assistants generally rely on explicit command triggering mechanisms, requiring users to deliberately interrupt ongoing communication, issue a specific wake word to the device, and then wait for a complete response. This forced, turn-based dialogue mode severely disrupts the continuity of natural human communication, making it difficult for voice assistants to truly integrate into multi-party conversational contexts. Furthermore, when the system detects the user's underlying intent within a continuous conversation flow, providing an answer using traditional voice broadcasting methods inevitably abruptly interrupts the user's ongoing voice channel.

[0003] Therefore, the technical problem to be solved in this application is how to analyze the semantics of user conversations in real time to capture users' implicit intentions in a multi-party conversational context and provide voice prompts through a companion response mechanism. Summary of the Invention

[0004] To address the above issues, this application provides an intelligent voice interaction method and system based on semantic recognition and call linkage, which greatly improves the efficiency of voice interaction.

[0005] To achieve the above objectives, the technical solution adopted in this application is as follows: In a first aspect, one specific embodiment of this application provides an intelligent voice interaction method based on semantic recognition and call linkage, comprising: The system acquires real-time audio streams of multi-party voice calls and performs temporal feature extraction and acoustic decoding. By combining context window annotations of sentence boundaries and semantic slots of historical rounds, a text sequence is obtained. Based on conditional random fields, context entity recognition and dependency extraction are performed on fuzzy word frequencies and intent categories in the text sequence to obtain semantic feature vectors carrying the semantic slots and triggering conditions; By comprehensively analyzing the confidence level corresponding to the intent category and the triggering conditions, the target information requirement corresponding to the semantic feature vector is determined. Based on the target information requirements, the target response text is obtained by retrieving field values ​​and definition entries that match the semantic slots from a pre-established structured knowledge base. The target response text is subjected to spectral envelope construction and resonant peak modulation to obtain a one-way prompt audio carrying harmonic distribution; The overlap rate between the one-way cue audio and the real-time audio stream is calculated, and the one-way cue audio is frequency-shifted by combining the masking value and phase difference to obtain the target cue audio; The spectral envelope of the target prompt audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis are calculated. After processing with a mixing and superposition method, the accompanying response audio is obtained and output.

[0006] Secondly, this application provides an intelligent voice interaction system based on semantic recognition and call linkage, comprising: The text sequence acquisition module is used to acquire the real-time audio stream of multi-party voice calls and perform temporal feature extraction and acoustic decoding. Combined with the context window to mark the sentence boundaries and the semantic slots of the historical rounds, the text sequence is obtained. The entity recognition module is used to perform context entity recognition and dependency extraction on the fuzzy word frequency and intent category in the text sequence based on the conditional random field, so as to obtain a semantic feature vector carrying the semantic slot and triggering condition; The analysis module is used to comprehensively analyze the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector. The retrieval and matching module is used to retrieve field values ​​and definition entries that match the semantic slots from a pre-established structured knowledge base according to the target information requirements, so as to obtain the target response text. The frequency domain analysis module is used to construct the spectral envelope and modulate the resonant peak position of the target response text to obtain a one-way prompt audio carrying harmonic distribution; The calculation module is used to calculate the overlap rate between the one-way prompt audio and the real-time audio stream, and to perform frequency band shifting processing on the one-way prompt audio by combining the masking value and phase difference to obtain the target prompt audio; The mixing and overlay output module is used to calculate the spectral envelope of the target cue audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis. After processing by mixing and overlay, the accompanying response audio is obtained and output.

[0007] This application provides an intelligent voice interaction method and system based on semantic recognition and call linkage. By extracting temporal features and acoustically decoding real-time audio streams, combining context window annotation of sentence boundaries and historical semantic slots, and then using conditional random fields to achieve entity recognition and dependency extraction from fuzzy word frequencies, it solves the problems of speech overlap, unclear sentence breaks, and ambiguous referents in multi-party calls, effectively improving the accuracy of extracting user intent and key information in colloquial expressions. Through comprehensive analysis of intent category confidence verification and trigger conditions, the target information requirement is determined, significantly reducing the probability of erroneous responses due to misidentification and improving the reliability of the interaction. In the speech synthesis stage, a prompt audio carrying harmonic distribution is constructed. Then, by matching the center frequency band, bandwidth, and harmonic overlap ratio, combined with masking thresholds and phase deviations, frequency band shifting is performed. Finally, accompanying response audio is generated through spectral envelope alignment and mixing superposition, achieving seamless integration of prompt sounds and call audio. This not only does not affect the user's call communication experience but also allows the target user to clearly receive the prompt information, achieving a seamless intelligent interaction loop in natural scenarios. Attached Figure Description

[0008] Figure 1 This is a schematic diagram illustrating the steps of an intelligent voice interaction method based on semantic recognition and call linkage in a specific embodiment of this application.

[0009] Figure 2 This is a schematic diagram of the system structure of an intelligent voice interaction method based on semantic recognition and call linkage in a specific embodiment of this application. Detailed Implementation

[0010] To enable those skilled in the art to better understand the technical solution, the present application will be described in detail below with reference to the embodiments. The description in this section is only exemplary and explanatory, and should not be used to limit the scope of protection of the present application in any way.

[0011] To address the shortcomings of existing voice interaction methods, such as reliance on wake-word triggers, breakup-based dialogue disrupting call continuity, interruption of voice communication due to response announcements, and inability to recognize implicit needs during calls, this application provides an intelligent voice interaction method and system based on semantic recognition and call linkage. By real-time parsing of multi-party call audio streams, combined with LSTM, temporal semantic annotation, and conditional random field entity extraction, it can accurately identify users' implicit information needs and complete ambiguous intentions based on historical confidence levels. Then, it retrieves and matches corresponding fields from a knowledge base and synthesizes prompt audio. Through frequency band shifting and phase compensation to optimize audio features, it finally achieves accompanying response output using a time-axis aligned mixing and overlay method. This eliminates the need for wake-word triggers and does not interfere with normal calls, achieving seamless intelligent interaction in natural scenarios.

[0012] Firstly, please refer to Figure 1As shown in a specific embodiment of this application, an intelligent voice interaction method based on semantic recognition and call linkage is provided, including the following steps: S1: Obtain the real-time audio stream of the multi-party voice call and perform temporal feature extraction and acoustic decoding. Combine the context window to mark the sentence boundaries and semantic slots of the historical rounds to obtain the text sequence. S2, Based on the conditional random field, perform context entity recognition and dependency extraction on the fuzzy word frequency and intent category in the text sequence to obtain a semantic feature vector carrying the semantic slot and triggering condition; S3, perform a comprehensive analysis of the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector; S4. Based on the target information requirements, retrieve the field values ​​and definition entries that match the semantic slots from the pre-established structured knowledge base to obtain the target response text; S5, construct the spectral envelope and modulate the resonant peak position of the target response text to obtain a one-way prompt audio carrying harmonic distribution; S6, calculate the overlap rate between the one-way prompt audio and the real-time audio stream, and perform frequency band shifting processing on the one-way prompt audio by combining the masking value and phase difference to obtain the target prompt audio; S7 calculates the spectral envelope of the target prompt audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis, processes them using a mixing and superposition method to obtain the accompanying response audio and output it.

[0013] This application provides an intelligent voice interaction method based on semantic recognition and call linkage. By extracting temporal features and acoustically decoding real-time audio streams, combining context window annotation of sentence boundaries and historical semantic slots, and then using conditional random fields to achieve entity recognition and dependency extraction from fuzzy word frequencies, it solves the problems of speech overlap, unclear sentence breaks, and ambiguous referents in multi-party calls, effectively improving the accuracy of extracting user intent and key information in colloquial expressions. Through comprehensive analysis of intent category confidence verification and trigger conditions, the target information requirement is determined, significantly reducing the probability of erroneous responses due to misidentification and improving the reliability of the interaction. In the speech synthesis stage, a prompt audio carrying harmonic distribution is constructed. Then, by matching the center frequency band, bandwidth, and harmonic overlap ratio, combined with masking thresholds and phase deviations, frequency band shifting is performed. Finally, accompanying response audio is generated through spectral envelope alignment and mixing superposition, achieving seamless integration of prompt sounds and call audio. This not only does not affect the user's call communication experience but also allows the target user to clearly receive the prompt information, achieving a seamless intelligent interaction loop in natural scenarios.

[0014] The steps in the embodiments provided in this application are described in detail below.

[0015] In step S1, the real-time audio stream of the multi-party voice call is acquired and subjected to temporal feature extraction and acoustic decoding. Combined with context window annotation of sentence boundaries and semantic slots of historical rounds, a text sequence is obtained, including the following steps: S11, acquire the real-time audio stream of the multi-party voice call, and perform silence removal processing on the real-time audio stream using an endpoint detection tool to obtain audio stream data; S12, based on the long short-term memory network, perform temporal feature extraction and acoustic decoding on the audio stream data to obtain the first acoustic frame sequence; S13, the first acoustic frame sequence is segmented according to a preset energy threshold to determine the set of punctuation points; S14, if the time interval between adjacent punctuation points in the punctuation point set is greater than a preset time threshold, then the first acoustic frame sequence between adjacent punctuation points is spliced ​​to obtain the second acoustic frame sequence. S15, obtain the semantic slot set corresponding to the historical rounds of dialogue, and synchronously annotate the second acoustic frame sequence and the semantic slot set according to the context window to obtain the annotation sequence; S16, the character conversion process of the annotation sequence is performed by a text conversion tool to obtain a text sequence.

[0016] It should be noted that the preset energy threshold refers to the short-time energy threshold used to determine the boundary between speech and silence. By calculating the short-time energy frame by frame in the first acoustic frame sequence, when the short-time energy of multiple consecutive frames is lower than the energy threshold, it is determined to be a silence / noise interval, and a sentence break point is marked at this point, thus achieving automatic sentence segmentation of the speech. This energy threshold is usually set based on the background noise level of the call scenario. For example, in a customer service scenario, the background noise is about -60dB, so the energy threshold can be set to -55dB.

[0017] The preset time threshold refers to the time limit for verifying the validity of sentence breaks, used to determine whether the speech segments between two adjacent sentence breaks belong to the same sentence. By calculating the time interval between adjacent sentence breaks, if the time interval is greater than the time threshold, it indicates that the sentence break within that interval is a misjudgment, such as a short pause by the user or a false break caused by environmental noise fluctuations. In this case, the two acoustic frame sequences need to be reassembled to eliminate invalid sentence breaks. If the time interval is less than or equal to the time threshold, it is considered a reasonable sentence break, and the original segmentation is retained. This time threshold can be set based on the normal pause duration of spoken expression, such as setting the time threshold to 1000ms.

[0018] Specifically, in a three-way customer service conference scenario, the communication server first receives mixed voice data from the customer, customer service, and technical experts via a real-time transmission protocol, and converts it into a real-time audio stream with a sampling rate of 16kHz and a bit depth of 16bit. An energy-based endpoint detection tool is used to calculate the short-time energy of the real-time audio stream frame by frame. When the energy of multiple consecutive frames is lower than a preset energy threshold (e.g., -55dB), it is identified as a silent segment and removed, resulting in the audio stream data. Then, the audio stream data is processed frame-by-frame, and 40-dimensional Mel-frequency cepstral coefficients are extracted as acoustic feature vectors. These vectors are then input sequentially into an LSTM (Long Short-Term Memory) network containing four hidden layers, each with 512 neurons. This LSTM network uses a forget gate to filter out irrelevant information such as ambient noise, retains key speech temporal dependencies through an input gate, and calculates the posterior probability of phonemes at each time step. Then, it combines a Hidden Markov Model (HMM) for acoustic decoding, that is, to find the maximum likelihood path in the search space, map the continuous audio frames into preliminary candidate Chinese character sequences, and finally obtain the first acoustic frame sequence.

[0019] It should be noted here that the first acoustic frame sequence refers to the frame-based composite sequence with timestamps obtained after processing the audio stream data (after silence removal) in frames, extracting temporal features, and performing acoustic decoding using a Long Short-Term Memory (LSTM) network. This composite sequence uses audio frames as the basic unit, with each frame synchronously carrying the original audio temporal waveform, frame-level acoustic features, phonemes / candidate Chinese characters obtained from acoustic decoding, and a global timestamp. It serves as an intermediate data carrier connecting audio signal processing and text semantic parsing. Short-time energy can be calculated from the intra-frame waveform data to achieve sentence segmentation detection, while temporal synchronous annotation of semantic slots and sentence segmentation markers is completed based on the decoded language and timestamps.

[0020] By calculating the short-time energy frame by frame in the first acoustic frame sequence, when the short-time energy of consecutive frames is detected to be lower than a preset energy threshold (e.g., -55dB) and the duration exceeds 200ms, the position is marked as a sentence break boundary point. The entire first acoustic frame sequence is traversed to obtain a sentence break point set composed of all sentence break boundary points. The sentence break point set is then traversed again, and the time interval between adjacent sentence break points is calculated. If the time interval is greater than a preset time threshold (e.g., 1000ms), it is determined to be an invalid sentence break. The first acoustic frame sequence within this interval is then reassembled into a continuous segment to eliminate the text fragmentation problem caused by incorrect sentence breakage, and finally the second acoustic frame sequence is obtained. Subsequently, the semantic slot set corresponding to the historical dialogue rounds is obtained, and a 500ms sliding context window is set. By backtracking the previous 3 historical dialogue rounds within this context window, the candidate Chinese characters decoded from the second acoustic frame sequence are pattern matched with the historical semantic slots. The words that match successfully are simultaneously labeled with the corresponding semantic slot labels, and punctuation marks are inserted at the sentence break boundaries to generate a label sequence carrying sentence break information and semantic slot labels. Finally, a text conversion tool is used to convert the phonemes and tag symbols in the tag sequence into standard Chinese characters, automatically complete the punctuation marks, and finally generate a text sequence with sentence boundary annotations and historical semantic slot annotations, providing structured input for subsequent semantic recognition.

[0021] The audio stream data is segmented into frames with a frame length of 25ms and a frame shift of 10ms. A Hamming window is added to each frame to suppress spectral leakage. 40-dimensional Mel-frequency cepstral coefficients are calculated for each audio frame, and first- and second-order difference coefficients are added to form a 40-dimensional acoustic feature vector sequence, which serves as the input to the LSTM network. Specifically, the LSTM network employs a four-layer unidirectional LSTM stacked structure, including an input layer for receiving the acoustic feature vector sequence arranged in time steps; four hidden layers, each containing 512 LSTM neurons, with fully connected layers transmitting temporal states. Each LSTM neuron in each layer includes a forget gate, an input gate, an output gate, and a cell state, using a gating mechanism to model the temporal dependencies of speech. Dropout layers are placed between LSTM neurons in each layer to prevent overfitting and enhance the model's generalization ability; the output layer uses a fully connected layer and a Softmax activation function, with an output dimension equal to the size of the modeled phoneme set (e.g., 600 dimensions), corresponding to the posterior probability distribution of phonemes at each time step. Then, a Hidden Markov Model (HMM) is used for acoustic decoding. The posterior probability of the phonemes is used as the emission probability of the HMM model. The Viterbi algorithm is used to find the maximum likelihood path in the search space. The phoneme sequence on the maximum likelihood path is converted into a candidate Chinese character sequence through the phoneme-to-Chinese character mapping. The candidate Chinese character sequence is bound with the timestamp information to form the first acoustic frame sequence, which provides temporally aligned text and acoustic information for subsequent sentence segmentation processing.

[0022] In step S2, context entity recognition and dependency extraction are performed on the fuzzy word frequency and intent category in the text sequence based on conditional random fields to obtain a semantic feature vector carrying the semantic slots and triggering conditions, including the following steps: S21, perform word segmentation on the text sequence, count the occurrences of each word, and determine the word frequency value; S22, if the word frequency value is greater than the preset word frequency threshold, then determine the intent category corresponding to the word; S23, Conditional Random Field is used to perform sequence labeling and syntactic analysis on the context corresponding to the intent category to identify entities and dependency relationships; S24, the identified entity and the dependency relationship are matched with the preset semantic slot template to obtain the semantic slot corresponding to the entity and the triggering condition corresponding to the dependency relationship; S25, the semantic slots and the triggering conditions are concatenated, and the intent category and original text features are fused to obtain a semantic feature vector carrying the semantic slots and triggering conditions.

[0023] It should be noted that the preset word frequency threshold refers to a pre-set word frequency statistics threshold used to distinguish between high-frequency ambiguous words and ordinary words in a text sequence. When the word frequency value of a word exceeds this word frequency threshold, the intent category determination process is triggered.

[0024] Preset semantic slot templates refer to structured templates pre-built for different intent categories, defining entity types, semantic slot names, and dependency parsing rules, used to map contextual entity recognition results into standardized semantic slots and triggering conditions.

[0025] Taking the user command text sequence "Help me adjust the living room air conditioner to a suitable temperature" as an example, the text sequence is first segmented using a word segmentation tool to obtain a vocabulary set: [help me, adjust, living room, air conditioner, adjust to, suitable, temperature]. Based on the historical corpus, the frequency of each word is counted. For example, if the frequency of the word "suitable" is determined to be 150 times, this frequency is compared with a preset frequency threshold (e.g., 100 times). Since 150 > 100, the word is determined to be a high-frequency fuzzy word. Then, based on the convolutional neural network intent classification model, the intent category corresponding to the text sequence is output as "device control" with a confidence level of 95%, providing intent prior for subsequent entity recognition.

[0026] Specifically, each word segmented is first mapped to a pre-trained word vector (such as Word2Vec / GloVe) to form a word vector matrix, which serves as the input to the intent classification model. The word vector for "appropriate" retains its semantic distribution information in the corpus, such as its high co-occurrence with words like "temperature," "brightness," and "volume." In the convolutional layer, convolutional kernels of different sizes slide across the word vector matrix to extract local n-gram features and form a feature map. For example, a 3-gram convolutional kernel captures the contextual fragment of "adjust to - appropriate - temperature" to strengthen the association between "appropriate" and "device adjustment-related words," while weakening the general modifying meaning. In the pooling layer, max pooling is performed on the feature map output from the convolutional layer, retaining the most significant signal in each feature channel, compressing the sequence length, and forming a fixed-length contextual feature vector. This condenses the association information between "appropriate" and surrounding words into a single key feature dimension. The pooled context feature vectors are then input into the fully connected layer, mapped to the intent category space, and the probability distribution of each intent category is output through the Softmax activation function. Based on the context features of "suitable," "temperature," "air conditioning," and "adjusted," the intent classification model assigns the highest probability to the "device control" category intent; for example, the probability of "device control" is output as 95%, while the probabilities of other intents are much lower than this value. The probability value output by the Softmax function is used as the confidence level. When the confidence level of "device control" reaches a preset confidence threshold (e.g., ≥90%), the intent category corresponding to "suitable" is determined to be "device control," with a confidence level of 95%.

[0027] Then, using the intent category of "device control" as a constraint, the text sequence is input into a conditional random field model. Combining part-of-speech tagging, vocabulary matching degree, and contextual features, the BIO annotation system is used to perform sequence annotation of words, identifying the location entity "living room" and the device entity "air conditioner". Through dependency parsing, a syntactic tree is constructed, determining that "living room" directly modifies "air conditioner", and the two have a noun-head modifier relationship, thus clarifying the semantic association between entities. Subsequently, the identified entities and dependency relations are matched with preset semantic slot templates (including location slots, device slots, and trigger condition slots), mapping the entity "living room" to the location slot and the entity "air conditioner" to the device slot. At the same time, the dependency relation "suitable temperature" is parsed as the temperature adjustment trigger condition, realizing a one-to-one correspondence between entities and semantic slots, and between dependency relations and trigger conditions. Finally, through vector concatenation technology, the state information of the location slots and device slots is fused with the trigger conditions and intent category features to generate a semantic feature vector.

[0028] In step S3, a comprehensive analysis is performed on the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector, including the following steps: S31, Obtain the initial text sequence corresponding to the intent category in the historical dialogue rounds, use a hidden Markov model to parse the initial text sequence, and calculate the confidence level corresponding to the intent category; S32, compare the confidence level with a preset confidence threshold range; S33, if the confidence level is within the preset confidence threshold range and the triggering condition is met, the target information requirement corresponding to the semantic feature vector is directly determined based on the maximum a posteriori probability estimation algorithm. S34, if the confidence level is outside the preset confidence threshold range, then perform a second comparison calculation on the fuzzy words based on the word frequency inverse document frequency algorithm, and supplement the missing semantic slots to obtain the completion item. Then, concatenate the completion item with the initial semantic vector corresponding to the completion item to obtain the completed target information requirement.

[0029] Specifically, taking a user's ambiguous refund request in an after-sales refund dialogue scenario as an example, the process begins by extracting initial text sequences related to the "after-sales refund" intent category from the first three historical dialogue rounds. These sequences are then used to construct the observation sequence for a Hidden Markov Model (HMM). The HMM uses the intent category as the hidden state and dialogue keywords as observation symbols. Based on pre-trained initial state probabilities, state transition probabilities, and emission probabilities, the Viterbi algorithm is used to solve for the most probable intent state transition path. During the calculation of the intent state transition path probability, the emission probabilities of keywords such as "refund" and "refund" are combined with the state transition probabilities between rounds for recursion. Finally, the intent state transition path probability is normalized, resulting in a confidence score of 0.82 for the "after-sales refund" intent category. The confidence score of 0.82 is compared with the preset confidence threshold range [0.75, 0.90]. It is confirmed that 0.82 falls within the confidence threshold range. At the same time, the triggering condition (such as the user having provided an order number) is detected to meet the requirements. At this time, based on the maximum a posteriori probability estimation algorithm, the confidence score of the current "after-sales refund" intent, the context information of the historical dialogue rounds, and the identified semantic slots (such as order number and product information) are used as input features. Combined with the pre-trained intent-demand mapping model, the posterior probability of the target information demand "execute a full refund" is calculated to be significantly higher than that of other alternative demands (such as "check refund progress" and "apply for exchange"). Therefore, the target information demand corresponding to this semantic feature vector is directly determined to be "execute a full refund", without the need to enter a secondary comparison process.

[0030] It should be noted that the pre-trained intent-demand mapping model refers to a classification model pre-trained based on historical dialogue data, used to establish the mapping relationship between user intent categories, semantic slot information, and target information demands. This intent-demand mapping model takes intent categories, confidence scores, semantic slot information, and contextual features as input. By learning the association rules between intents and business actions in historical dialogues, it outputs the posterior probability of each target information demand. Based on the maximum posterior probability, it determines the final target information demand, achieving accurate conversion from natural language intents to executable business actions.

[0031] If the calculated confidence score is 0.65, falling outside the preset confidence threshold range [0.75, 0.90], a backtracking mechanism is initiated to return to the context window. Unmatched character sequences are extracted and segmented to obtain vague words such as "refund," "money," and "not received." Then, the inverse document frequency (IVF) algorithm is used to calculate the frequency weight of each vague word in the context. For example, the local frequency weight of the character "refund" is calculated to be as high as 0.45. This frequency weight is then compared a second time with the baseline frequency in the standard lexicon to select candidate words that match the vague word frequency features and construct completion items to fill in the missing semantic slots. Finally, the completion items are concatenated with the corresponding initial semantic vectors, and the user intent evolution trend is deduced using the state transition probability matrix of a Markov chain model. The completed target information requirement is then determined to be "query refund progress."

[0032] In step S4, based on the target information requirement, the field values ​​and definitions matching the semantic slots are retrieved from a pre-established structured knowledge base to obtain the target response text, including the following steps: S41, perform entity recognition on the target information requirement to obtain the first semantic slot set corresponding to the entity; S42, by traversing the pre-established structured knowledge base, retrieve the knowledge entries corresponding to the first semantic slot set to form a candidate set; S43, calculate the similarity between the first semantic slot set and the candidate set using the cosine similarity algorithm to obtain a similarity set; S44, if any similarity in the similarity set is greater than a preset similarity threshold, then the target field value is determined based on the candidate set corresponding to the current similarity. S45, query the definition text corresponding to the target field value in the pre-established structured knowledge base, and obtain the target response text after performing part-of-speech tagging on the definition text.

[0033] It should be noted that the pre-built structured knowledge base is a collection of business knowledge organized in a field-based and item-based manner. It is usually stored in the form of database tables or JSON objects. Each knowledge entry contains slot matching conditions, such as business type, query object, status value, target field value and explanatory text, which are used to provide corresponding response content for different intents and slot combinations.

[0034] The preset similarity threshold is a pre-defined cosine similarity threshold used to determine whether the matching degree between candidate knowledge items in the structured knowledge base and the user's semantic slot set meets the usable standard. It is typically set according to the fault tolerance of the business scenario. For example, in an after-sales refund scenario, if the preset similarity threshold is set to 0.85, and the similarity between the knowledge items "Refund Progress Query" and "Refund in Progress" is 0.92 (>0.85), then it is considered a successful match.

[0035] The definition text is the standardized natural language text in the structured knowledge base that corresponds to the value of the target field. The final response content generated according to the user's intent can be directly output to the user after part-of-speech tagging and sentence optimization.

[0036] Let's take the scenario of "querying refund progress" as an example. First, entity recognition is performed on the target information requirement "querying refund progress," extracting the core intent slot as "refund progress query," and associating it with a semantic slot set: {intent type: after-sales query, business type: refund, query object: progress}, forming a structured first semantic slot set, which serves as the query condition for knowledge base retrieval. Then, the pre-established structured knowledge base is traversed. In this embodiment, it can be an after-sales business structured knowledge base. Based on the slot combination of "after-sales query - refund - progress," all knowledge entries containing fields related to "refund progress" are filtered out to obtain a candidate set, such as knowledge entries containing the field values ​​"refund in progress," "received," and "processing." Next, the first semantic slot set is converted into a feature vector, and cosine similarity is calculated between it and the feature vectors of each knowledge entry in the candidate set. For example, the similarity score for the knowledge item "Refund progress query - Refund in progress" is calculated to be 0.92; for "Refund progress query - Funds received" it is calculated to be 0.78; and for "Refund progress query - Processing" it is calculated to be 0.85. A similarity set containing the similarities of each knowledge item is generated. The values ​​in the similarity set are then compared with a preset similarity threshold (e.g., 0.85). The similarity score of "Refund in progress" (0.92) is greater than the preset threshold, indicating a successful match. Therefore, the corresponding target field value is determined to be "Refund in progress". Finally, the definition text corresponding to the field value "Refund in progress" is queried in the structured knowledge base. The definition text is then subjected to part-of-speech tagging and sentence optimization to remove redundant expressions, resulting in the target response text, such as "Your refund application is being processed and is expected to arrive within 3 business days. Please pay attention to account changes." This provides accurate text input for subsequent speech synthesis.

[0037] In step S5, the target response text is subjected to spectral envelope construction and resonant peak modulation to obtain a unidirectional cue audio carrying harmonic distribution, including the following steps: S51, perform acoustic conversion processing on the target response text to obtain an initial speech stream; S52, calculate the amplitude spectrum of each frame of the initial speech stream based on the Fast Fourier Transform algorithm, and connect the peak points of each formant in the amplitude spectrum in sequence to form the spectral envelope of the corresponding frame. S53, perform linear prediction coding analysis on the spectral envelope, and obtain the corresponding pole positions by solving the linear prediction coefficients; S54, According to the modulation target, the frequency and bandwidth of the resonance peak are modulated by adjusting the position of the poles to obtain the modulated spectrum parameters; S55 processes the modulated spectral parameters based on the encoder and parses the preset fundamental frequency trajectory to generate an excitation signal; S56, the excitation signal is restored to a time-domain waveform to generate a unidirectional prompt audio carrying harmonic distribution.

[0038] Let's take the target response text "Your refund application is being processed and is expected to arrive within 3 business days. Please pay attention to account changes" in an after-sales refund scenario as an example. First, the target response text undergoes acoustic conversion processing (such as speech synthesis technology), that is, it is converted into an initial speech stream at a sampling rate of 16kHz and a bit depth of 16bit. This initial speech stream is in mono PCM format and contains complete sentence pronunciation, pauses, and intonation information. Then, the initial speech stream is processed into frames, with a frame length typically set to 20-30ms (e.g., 25ms), and a frame shift of 10ms to ensure partial overlap between adjacent frames and avoid spectral leakage. A Hamming window is then added to each frame to weaken the discontinuity at frame edges. A Fast Fourier Transform is performed on the single frame of speech after the Hamming window is applied, converting the discrete time-domain signal into a discrete frequency-domain signal. By calculating the magnitude of this discrete frequency-domain signal, the amplitude spectrum, also known as the power spectrum, is obtained, which refers to the energy distribution of the speech signal at different frequencies. The horizontal axis of the amplitude spectrum represents frequency, ranging from 0 to Fs / 2, where Fs is the sampling rate (e.g., 0-8kHz for 16kHz). The vertical axis represents amplitude, i.e., energy level. The amplitude spectrum contains information about the fundamental frequency, harmonics, and formants of speech. Extracting the peak points of the formants in the amplitude spectrum and connecting these peak points forms a continuous spectral envelope, reflecting the resonance characteristics of the vocal tract.

[0039] Next, linear predictive coding (LPC) analysis is performed on the spectral envelope. The speech signal is treated as the output of an all-pole filter excited by white noise or a pulse sequence. The transfer function of the vocal tract is estimated by solving for the linear prediction coefficients (i.e., LPC coefficients). Then, polynomial roots are taken from the LPC coefficients to obtain the corresponding pole positions, which correspond to the frequency and bandwidth of the formants. For example, through LPC analysis, the first formant F1 is identified as approximately 500Hz, and the second formant F2 as approximately 1500Hz. Subsequently, the target frequency offset of the formants is determined according to the modulation target. By moving the pole positions on the complex plane, the frequency and bandwidth of the formants are changed, achieving formant position modulation. For example, to make the timbre sound softer, the first formant F1 can be moved to a lower frequency, and the second formant F2 can be moved to a higher frequency. For frequency modulation, the center frequency of the resonant peak is shifted by changing the angle (phase) of the poles. For bandwidth modulation, it is achieved by changing the distance between the poles and the origin. The closer the distance, the narrower the bandwidth and the sharper the resonant peak; the farther the distance, the wider the bandwidth and the flatter the resonant peak. For example, by moving the pole of the first resonant peak F1 from 500Hz to 450Hz and the pole of the second resonant peak F2 from 1500Hz to 1600Hz, while keeping the distance between the poles and the origin constant and only adjusting their angles, resonant peak position modulation can be achieved.

[0040] The modulated pole positions are converted back to LPC coefficients to obtain a new set of linear prediction coefficients. These, along with the corresponding fundamental frequency, gain, and other information, constitute the modulated spectral parameters, fully describing the modulated channel characteristics. The encoder then processes the modulated spectral parameters and analyzes the preset fundamental frequency trajectory to generate an excitation signal. This excitation signal is input into a linear prediction synthesis filter constructed from the spectral parameters for filtering, reconstructing a speech time-domain waveform containing harmonic distributions. After smoothing, noise reduction, and volume normalization, a unidirectional prompt audio carrying harmonic distributions is finally generated, which can be directly played to the user.

[0041] In step S55, the modulated spectral parameters are processed based on the encoder, and the preset fundamental frequency trajectory is parsed to generate an excitation signal, including the following steps: S551 converts modulated spectral parameters into channel filter parameters based on an encoder; S552 analyzes the preset baseband trajectory and generates a frame excitation signal corresponding to the baseband trajectory for each frame of speech. S553, the channel filter parameters are associated with the frame excitation signal to obtain the excitation signal.

[0042] It should be noted that the preset fundamental frequency trajectory refers to a pre-set fundamental frequency value sequence arranged according to the time sequence of speech frames. It is used to define the pitch and intonation variation patterns of synthesized speech at different time periods, distinguish between voiced and unvoiced speech intervals, and provide frequency basis for the generation of excitation signals.

[0043] The encoder converts the modulated spectral parameters (LPC coefficients) into tract filter parameters. If the input is LPC coefficients, an all-pole linear prediction filter is directly constructed; if the input is pole locations, the poles are mapped back to LPC coefficients through the inverse process of polynomial root finding to construct the tract transfer function. A preset fundamental frequency trajectory is then analyzed. This trajectory defines the fundamental frequency (F0) value of each speech frame in time-series form, for example, a mean of 120Hz and a fluctuation range of ±20Hz. It also includes voiced / unvoiced markers. Voiced frames have a non-zero fundamental frequency, corresponding to the voiced portion of vocal cord vibration; unvoiced frames have a fundamental frequency of 0, corresponding to unvoiced portions such as fricatives and plosives. The encoder maps the fundamental frequency trajectory to each speech frame, generating a corresponding excitation source type for each frame. For example, voiced frame excitation generates a pulse sequence with the fundamental frequency as the period, and the pulse interval is determined by the fundamental frequency; the lower the fundamental frequency, the larger the pulse interval. Unvoiced frame excitation generates a Gaussian white noise sequence to simulate friction / pop noise without a fundamental frequency. Mixed frame processing: for speech segments with a mixture of voiced and unvoiced sounds, an excitation signal is used that is a proportional mixture of pulse sequences and white noise to preserve the details of natural speech. The encoder correlates the generated frame-by-frame excitation signals with the vocal tract filter parameters to obtain the final excitation signal. That is, the encoder adds gain control to each frame excitation signal to match the loudness of the target response text and smoothly transitions the excitation signals of adjacent frames to avoid speech distortion caused by abrupt changes between frames. Finally, a continuous excitation signal sequence is generated that perfectly matches the modulated spectral parameters and fundamental frequency trajectory.

[0044] In step S6, the overlap rate between the one-way cue audio and the real-time audio stream is calculated, and the one-way cue audio is frequency-shifted by combining the masking value and phase difference to obtain the target cue audio, including the following steps: S61, Based on the Fast Fourier Transform algorithm, the one-way prompt audio and the real-time audio stream are converted in the time-frequency domain to obtain the first spectrum matrix and the second spectrum matrix; S62, extract the mid-frequency band of the one-way prompt audio and the bandwidth of the real-time audio stream; S63, perform harmonic group analysis on the first spectrum matrix and the second spectrum matrix using the mid-frequency band and the bandwidth to determine the overlap rate between the one-way prompt audio and the real-time audio stream; S64, if the overlap rate is greater than the preset overlap value, then calculate the auditory masking effect on the first spectrum matrix and the second spectrum matrix to obtain the masking value and phase difference; S65, perform frequency band shifting on the mid-frequency band based on the masking value and the phase difference to obtain target prompt audio with limited shift amplitude.

[0045] It should be noted that the preset overlap value refers to the pre-set threshold for the spectral energy overlap rate, which is used to determine whether the degree of spectral overlap between the one-way prompt audio and the real-time audio stream in the key frequency band reaches the standard that requires frequency band optimization. When the overlap rate is greater than the overlap value, auditory masking calculation and frequency band shifting are triggered.

[0046] Limited translation amplitude refers to the constraint imposed on the translation operation of the one-way prompt audio frequency band. By setting a preset maximum translation threshold, the amplitude of the frequency band shift is controlled to avoid distortion and decreased clarity of the one-way prompt audio due to excessive adjustment. Specifically, the preset maximum translation threshold is a pre-defined upper limit for the frequency band shift, used to control the translation amplitude of the mid-frequency band in the one-way prompt audio. This prevents damage to the speech formant distribution and harmonic structure due to excessive adjustment, thereby ensuring the clarity and intelligibility of the target prompt audio while resolving spectral overlap interference.

[0047] This paper describes the processing of a one-way prompt audio message ("Your refund application is being processed and is expected to arrive within 3 business days") generated in a post-sales refund scenario, along with a real-time audio stream. First, the one-way prompt audio and the real-time audio stream are processed by frame-by-frame windowing at a sampling rate of 16kHz, a frame length of 25ms, and a frame shift of 10ms. A Fast Fourier Transform is then performed on each frame to obtain the first spectrum matrix (frequency domain amplitude and phase distribution of the one-way prompt audio) and the second spectrum matrix (frequency domain amplitude and phase distribution of the real-time audio stream). Rows in the matrices correspond to time frames, and columns correspond to frequency points. Next, the mid-frequency bands with concentrated energy in the one-way prompt audio, such as 500Hz-3000Hz, are extracted, covering the fundamental frequency and main formants of the speech. The bandwidth of the real-time audio stream is calculated; for example, the effective frequency range of the current user's speech is 200Hz-3500Hz, providing frequency boundaries for subsequent harmonic analysis.

[0048] Then, constrained by the mid-frequency band and bandwidth, harmonic group analysis is performed on the first and second spectral matrices. This involves comparing the harmonic distribution of the two spectral matrices in the mid-frequency band. For example, significant harmonic peaks are found at 1000Hz and 1500Hz for the one-way audio prompt, while energy peaks are found at 1100Hz and 1600Hz for the real-time audio stream. Calculations show that the frequency points where the energy overlaps between the first and second spectral matrices in the mid-frequency band is 72%, i.e., an overlap rate of 0.72. Comparing this overlap rate of 0.72 with a preset overlap value (e.g., 0.6), since 0.72 > 0.6, frequency band optimization is determined. At this point, the auditory masking effect is calculated based on the two spectral matrices, resulting in a masking value of -12dB for the one-way audio prompt by the real-time audio stream, meaning that the one-way audio prompt is masked by the user's voice energy at some frequencies. Simultaneously, the phase difference between the real-time audio stream and the one-way audio prompt in the mid-frequency band is calculated, with the average phase difference at key peak points set at 45° to reflect the degree of phase interference between the signals.

[0049] Finally, based on the masking value of -12dB and a phase difference of 45°, the mid-frequency band of the one-way prompt audio is shifted by 150Hz towards the higher frequencies. This shifts the harmonic peaks of the one-way prompt audio away from the energy concentration area of ​​the real-time audio stream, such as 1000Hz→1150Hz and 1500Hz→1650Hz. Simultaneously, the signal phase is adjusted according to the phase difference to reduce superposition interference in the frequency domain. The result is a target prompt audio that ensures clarity while avoiding conflict with the user's speech. In other words, the frequency shifting process in step S6 operates in the frequency domain, shifting the mid-frequency band of the one-way prompt audio based on the spectral overlap rate, masking value, and phase difference. The aim is to reduce spectral interference between the one-way prompt audio and the real-time audio stream, ensuring speech clarity.

[0050] In step S7, the spectral envelope of the target cue audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis are calculated. The accompanying response audio is obtained and output using a mixing and superposition method, including the following steps: S71, The target prompt audio and the real-time audio stream are processed based on the Fast Fourier Transform algorithm to obtain the first amplitude spectrum and the second amplitude spectrum; S72, extract the spectral envelope of the first amplitude spectrum and the harmonic distribution of the second amplitude spectrum to obtain the initial acoustic features; S73, Based on the initial acoustic characteristics, calculate the alignment offset of the spectral envelope and the harmonic distribution on the time axis using a dynamic time warping algorithm; S74, compare the alignment offset with a preset offset threshold, and combine it with translation amplitude compensation to obtain the compensated acoustic features; S75, the compensated acoustic features are mixed and superimposed using a mixing track to obtain the accompanying response audio; S76, the accompanying response audio is injected into the voice channel output based on the audio transmission protocol.

[0051] It should be noted that the preset offset threshold is a time value (e.g., 0.2 seconds), representing the maximum acceptable timing misalignment during mixing. Exceeding this offset threshold will result in a noticeable desynchronization between the target audio cue and the real-time audio stream. By determining the degree of timing misalignment between the target audio cue and the real-time audio stream using the preset offset threshold, it is determined whether translational compensation of the target audio cue is needed. The final output is a compensated acoustic feature that can be directly used for mixing.

[0052] The processing is described using the target prompt audio (i.e., the refund progress prompt audio after frequency band shifting optimization) and the real-time audio stream (the user's refund inquiry voice) as the processing objects. First, the target prompt audio and the real-time audio stream are processed by frame-by-frame windowing at a sampling rate of 16kHz, a frame length of 25ms, and a frame shift of 10ms. Then, a Fast Fourier Transform is performed to obtain the first amplitude spectrum of the target prompt audio and the second amplitude spectrum of the real-time audio stream. Both amplitude spectra are two-dimensional matrices with rows corresponding to time frames and columns corresponding to frequency points. Next, the energy profile is extracted from the first amplitude spectrum to obtain the spectral envelope of the target prompt audio, reflecting the formants and timbre characteristics of the target prompt audio; and the peak distribution of each harmonic of the real-time audio stream is extracted from the second amplitude spectrum to obtain the harmonic distribution of the real-time audio stream. The spectral envelope and harmonic distribution are combined to form the initial acoustic features.

[0053] Then, using the initial acoustic features as input, the Dynamic Time Warping (DTW) algorithm is employed to calculate the optimal matching path on the time axis between the spectral envelope of the target cue audio and the harmonic distribution of the real-time audio stream, obtaining the alignment offset between the two signals. Specifically, feature sequences are constructed for the spectral envelope of the target cue audio and the harmonic distribution of the real-time audio stream, respectively, according to time frames, resulting in the cue audio sequence and the real-time audio sequence. A distance matrix is ​​then constructed by calculating the distance between each pair of feature vectors in the cue audio sequence and the real-time audio sequence. The distance can be Euclidean or cosine distance, and the elements in the distance matrix represent the feature difference between the i-th frame of the cue audio sequence and the j-th frame of the real-time audio sequence, where the i-th frame and the j-th frame correspond as a pair of feature vectors. The cumulative distance matrix is ​​initialized, and the minimum cumulative distance path is solved recursively using dynamic programming. After the recursion is completed, the optimal warped path connecting the cue audio sequence and the real-time audio sequence is obtained by backtracking from the end point of the matrix. This optimal warped path reflects the optimal time correspondence between the two signal frames of the target cue audio and the real-time audio stream. Based on the overall time misalignment of the target prompt audio and the real-time audio stream using the optimal warping path, if the optimal warping path is biased towards the beginning of the real-time audio stream, it indicates that the target prompt audio is lagging and needs to be played earlier; if the optimal warping path is biased towards the beginning of the target prompt audio, it indicates that the target prompt audio is ahead and needs to be played later. The final calculated time difference is the alignment offset required for subsequent mixing. That is, the translation amplitude compensation processing in step S7 acts on the time domain, and the target prompt audio is time-aligned based on the timing offset obtained by the dynamic time warping algorithm, with the aim of eliminating playback timing deviations.

[0054] The calculated alignment offset is compared with a preset offset threshold. If the alignment offset is greater than the preset threshold, it indicates that the timing misalignment between the target prompt audio and the real-time audio stream is too large, requiring translation amplitude compensation. This involves shifting the spectral envelope of the target prompt audio along the time axis based on the magnitude and direction (advanced or delayed) of the alignment offset. Simultaneously, to avoid audio dropouts or abruptness caused by the shift, the first and last frames of the target prompt audio are faded in and out. The spectral envelope after translation amplitude compensation, along with the harmonic distribution of the real-time audio stream, together constitute the compensated acoustic features. If the alignment offset is less than or equal to the preset offset threshold, it indicates that the timing misalignment between the target prompt audio and the real-time audio stream is within an acceptable range, requiring no significant adjustment. In this case, the spectral envelope of the target prompt audio and the harmonic distribution of the real-time audio stream are directly extracted as the compensated acoustic features, maintaining the original signal timing unchanged. Regardless of whether translation amplitude compensation is applied, the final output compensated acoustic features include the spectral envelope of the target cue audio after time alignment and the harmonic distribution of the real-time audio stream, providing time-matched acoustic data for subsequent mixing and overlay.

[0055] Finally, the time-aligned and translationally compensated target cue audio is mixed and superimposed with the real-time audio stream through a mixing track. That is, the gain ratio of the target cue audio and the real-time audio stream is adjusted according to the acoustic characteristics after translational compensation to ensure that the target cue audio is clearly audible and does not drown out the human voice. Finally, the accompanying response audio is generated and output to the playback device.

[0056] In step S76, injecting the accompanying response audio into the voice channel output based on the audio transmission protocol includes the following steps: S761, Based on the Hidden Markov Model, the accompanying response audio is parsed to obtain the audio state sequence; S762, The audio state sequence is classified using a support vector machine to determine the audio category label; S763, if the audio category label meets the preset category conditions, then the accompanying response audio is resampled according to the audio transmission protocol to obtain the first accompanying response audio; S764, Encode the first accompanying response audio to obtain an audio protocol data packet; S765, obtain the current buffer pool capacity of the target voice channel; if the current buffer pool capacity is greater than a preset capacity threshold, inject the audio protocol data packet into the target voice channel and obtain the injection status code. S766, if the injected status code matches the preset success status identifier, then generate and output the confirmation command for the interactive closed loop.

[0057] In step S76, it should be noted that the preset category conditions refer to the pre-set audio category admission standards, which specify the range of audio category labels allowed to enter the transmission process. They are used to determine whether the current audio complies with the transmission protocol requirements and serve as the basis for compliance determination before audio transmission.

[0058] The preset capacity threshold refers to the pre-set critical value of the available capacity of the target voice channel buffer pool. It is used to determine whether the target voice channel has the conditions to receive audio protocol data packets. When the current capacity of the buffer pool is greater than the capacity threshold, the injection of audio protocol data packets is allowed.

[0059] The preset success status identifier refers to a predefined status code or identifier that represents the successful injection of audio protocol data packets. It is used to compare with the actual injection status code to confirm whether the audio protocol data packets have been successfully injected into the target voice channel.

[0060] In one specific embodiment, the accompanying response audio is first divided into fixed-length frames (e.g., 25ms per frame). Quantifiable acoustic features are extracted frame by frame, including frame energy (average signal strength of the frame); zero-crossing rate (number of times the frame crosses zero); fundamental frequency estimation (calculated directly using autocorrelation or cepstral method); and duration status marking (marking the frame as a "sound segment" or "silent segment" based on the energy and zero-crossing rate of consecutive frames). Then, the audio features of each frame are directly determined based on preset rules to classify the audio state. The preset rules may include: if the frame energy is higher than the preset energy threshold, the zero-crossing rate is lower than the preset unvoiced threshold, and the fundamental frequency is within the normal human voice range (e.g., 80-400Hz), then it is marked as a voiced state; if the frame energy is higher than the preset energy threshold, the zero-crossing rate is higher than the preset unvoiced threshold, and the fundamental frequency is 0, then it is marked as an unvoiced state; if the frame energy is lower than the preset energy threshold, then it is marked as a silent state; the state marks are arranged sequentially according to the frame order to form a complete audio state sequence, for example: [voiced, voiced, unvoiced, voiced, silent].

[0061] It should be noted here that the preset rules refer to the predefined conditional judgment logic and fixed thresholds based on quantified acoustic features, which are used to directly determine the business category of the state and sequence of audio frames without relying on machine learning model training and prediction. They have the characteristics of logical transparency and strong interpretability.

[0062] The preset energy threshold refers to the pre-set critical value of audio signal energy, which is used to determine whether there is a valid speech signal in the audio frame by frame, thereby distinguishing between sound frames and silent frames. It is the basic judgment parameter for audio state analysis.

[0063] The unvoiced threshold is a pre-set zero-crossing rate threshold used to distinguish between unvoiced and voiced sounds based on the magnitude of the zero-crossing rate of the audio frame. It is a parameter for determining the audio state.

[0064] Then, the audio state sequence is traversed, and key quantitative indicators of the audio state sequence are statistically analyzed, including the duration and proportion of each state, such as voiced sound proportion of 70%, unvoiced sound proportion of 20%, and silence proportion of 10%; total sequence duration, number of state transitions, and maximum duration of continuous silence segments; and the average value and fluctuation range of the fundamental frequency of voiced segments. The statistical indicators are matched and judged according to preset category rules: if the total sequence duration is within the preset prompt tone duration range, the voiced sound proportion is ≥60%, the silence segment is ≤500ms, and the fundamental frequency fluctuation is within the preset range, it is judged as the "standard prompt tone" category; if the silence segment of the sequence is too long and the voiced sound proportion is extremely low, it is judged as the "invalid noise" category; if the fundamental frequency fluctuation range and state distribution conform to human voice characteristics, it is judged as the "user voice" category; after a successful match, the corresponding audio category label is directly output.

[0065] It should be noted here that the preset category rules refer to multiple predefined judgment conditions, which are matched with statistical parameters such as the duration of the audio state sequence, the proportion of each state, the silence duration, and the fundamental frequency range, in order to distinguish the audio service type and determine the category label.

[0066] The preset prompt tone duration range refers to the pre-defined total duration range of standard prompt audio, which is used to verify whether the actual duration of the current audio meets the specification requirements and is a basic condition for determining the audio category.

[0067] The preset range refers to a predefined range of parameter values, including upper and lower limits, used to compare measured parameters and determine whether they are within the specified range, serving as the basis for status determination and category differentiation.

[0068] Next, by verifying whether the audio category label "Customer Service Standard Prompt Tone" meets the preset category conditions (such as sampling rate and format compliance), and confirming that the preset category conditions are met, the accompanying response audio is resampled from 16kHz to the 8kHz format specified by the audio transmission protocol (such as RTP protocol) to obtain the first accompanying response audio, making it conform to the standard sampling rate requirements of the transmission link. Then, the 8kHz first accompanying response audio is encoded using an encoder (such as a G.711 encoder) to convert the PCM format audio data into a bitstream conforming to the audio transmission protocol standard, and fragmented and encapsulated at 20ms / packet to generate audio protocol data packets containing timestamps, sequence numbers, and check information, ensuring the integrity and recoverability of data transmission.

[0069] Next, the current buffer pool capacity of the target voice channel is obtained. If the current buffer pool capacity is greater than the preset capacity threshold (e.g., the current buffer pool capacity is 80% and the preset capacity threshold is 50%), it indicates that the target voice channel has receiving capability. Then, audio protocol data packets are injected into the target voice channel sequentially, and an injection status code is obtained. For example, an injection status code of "0x00" indicates successful injection. The injection status code "0x00" is then compared with the preset success status identifier "0x00" to confirm that the injection status code matches the preset success status identifier. Subsequently, an interactive closed-loop confirmation command is generated, such as "[Interactive Closed-Loop Confirmation] Refund prompt tone has been successfully sent to the voice channel," completing this voice interaction process.

[0070] Please see Figure 2 As shown, in a second aspect, this application provides an intelligent voice interaction system based on semantic recognition and call linkage, comprising: The text sequence acquisition module is used to acquire the real-time audio stream of multi-party voice calls and perform temporal feature extraction and acoustic decoding. Combined with the context window to mark the sentence boundaries and the semantic slots of the historical rounds, the text sequence is obtained. The entity recognition module is used to perform context entity recognition and dependency extraction on the fuzzy word frequency and intent category in the text sequence based on the conditional random field, so as to obtain a semantic feature vector carrying the semantic slot and triggering condition; The analysis module is used to comprehensively analyze the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector. The retrieval and matching module is used to retrieve field values ​​and definition entries that match the semantic slots from a pre-established structured knowledge base according to the target information requirements, so as to obtain the target response text. The frequency domain analysis module is used to construct the spectral envelope and modulate the resonant peak position of the target response text to obtain a one-way prompt audio carrying harmonic distribution; The calculation module is used to calculate the overlap rate between the one-way prompt audio and the real-time audio stream, and to perform frequency band shifting processing on the one-way prompt audio by combining the masking value and phase difference to obtain the target prompt audio; The mixing and overlay output module is used to calculate the spectral envelope of the target cue audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis. After processing by mixing and overlay, the accompanying response audio is obtained and output.

[0071] This application, by introducing speech recognition technology, constructs a complete natural interaction loop, breaking through the limitations of traditional one-way audio prompts. It achieves real-time perception and understanding of user intent, significantly improving the naturalness of voice interaction and user experience. On one hand, by capturing user voice input in real time, it avoids interference between prompts and user speech, reducing unnecessary interruptions and repetitive playback, making the dialogue flow closer to real-life communication. On the other hand, the recognition results can be linked with business logic, enabling multi-turn dialogue context understanding and scene-adaptive responses. It can dynamically adjust the content, speed, and tone of accompanying response audio according to user intent, improving business processing efficiency and service adaptability.

[0072] It should be noted that, in this document, the terms "comprising," "including," and any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Specific examples have been used in this document to illustrate the principles and implementation methods of the technical solutions of this application. The above examples are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are merely preferred embodiments of this application. It should be pointed out that, due to the limitations of written expression and the objective existence of infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this application, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes, or combinations, or the direct application of the concept and technical solutions of this application to other situations without modification, should all be considered within the scope of protection of this application.

Claims

1. An intelligent voice interaction method based on semantic recognition and call linkage, characterized in that, include: The system acquires real-time audio streams of multi-party voice calls and performs temporal feature extraction and acoustic decoding. By combining context window annotations of sentence boundaries and semantic slots of historical rounds, a text sequence is obtained. Based on conditional random fields, context entity recognition and dependency extraction are performed on the fuzzy word frequency and intent category in the text sequence to obtain a semantic feature vector carrying the semantic slot and triggering condition; By comprehensively analyzing the confidence level corresponding to the intent category and the triggering conditions, the target information requirement corresponding to the semantic feature vector is determined. Based on the target information requirements, the target response text is obtained by retrieving field values ​​and definition entries that match the semantic slots from a pre-established structured knowledge base. The target response text is subjected to spectral envelope construction and resonant peak modulation to obtain a one-way prompt audio carrying harmonic distribution; The overlap rate between the one-way cue audio and the real-time audio stream is calculated, and the one-way cue audio is frequency-shifted by combining the masking value and phase difference to obtain the target cue audio; The spectral envelope of the target prompt audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis are calculated. After processing with a mixing and superposition method, the accompanying response audio is obtained and output.

2. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, The system acquires real-time audio streams of multi-party voice calls, performs temporal feature extraction and acoustic decoding, and combines context window annotations of sentence boundaries and semantic slots of historical rounds to obtain a text sequence, including: The real-time audio stream of a multi-party voice call is acquired, and the real-time audio stream is muted using an endpoint detection tool to obtain audio stream data. The audio stream data is processed by extracting temporal features and performing acoustic decoding based on a long short-term memory network to obtain the first acoustic frame sequence. The first acoustic frame sequence is segmented according to a preset energy threshold to determine the set of punctuation points; If the time interval between adjacent punctuation points in the set of punctuation points is greater than a preset time threshold, then the first acoustic frame sequence between adjacent punctuation points is spliced ​​together to obtain the second acoustic frame sequence. Obtain the semantic slot set corresponding to the historical rounds of dialogue, and synchronously annotate the second acoustic frame sequence and the semantic slot set according to the context window to obtain the annotation sequence; The annotation sequence is processed by a text conversion tool to obtain a text sequence.

3. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, Based on conditional random fields, contextual entity recognition and dependency extraction are performed on the fuzzy word frequency and intent category in the text sequence to obtain a semantic feature vector carrying the semantic slots and triggering conditions, including: The text sequence is segmented into words, and the frequency of each word is counted to determine the word frequency value; If the word frequency value is greater than the preset word frequency threshold, then the intent category corresponding to the word is determined; Conditional random fields are used to perform sequence labeling and syntactic analysis on the context corresponding to the intent category to identify entities and dependency relationships; The identified entities and dependencies are matched with preset semantic slot templates to obtain the semantic slots corresponding to the entities and the triggering conditions corresponding to the dependencies. By concatenating the semantic slots and the triggering conditions, and fusing the intent category with the original text features, a semantic feature vector carrying the semantic slots and triggering conditions is obtained.

4. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, A comprehensive analysis is performed on the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector, including: Obtain the initial text sequence corresponding to the intent category in the historical dialogue rounds, parse the initial text sequence using a Hidden Markov Model, and calculate the confidence level corresponding to the intent category; The confidence level is compared with a preset confidence threshold range; If the confidence level is within the preset confidence threshold range and the triggering condition is met, the target information requirement corresponding to the semantic feature vector is directly determined based on the maximum a posteriori probability estimation algorithm. If the confidence level is outside the preset confidence threshold range, the fuzzy words are compared and calculated a second time based on the word frequency inverse document frequency algorithm, and the missing semantic slots are supplemented to obtain the completion item. The completion item is then concatenated with the initial semantic vector corresponding to the completion item to obtain the completed target information requirement.

5. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, Based on the target information requirements, the target response text is obtained by retrieving field values ​​and definitions that match the semantic slots from a pre-established structured knowledge base, including: Entity recognition is performed on the target information requirement to obtain the first semantic slot set corresponding to the entity; By traversing the pre-established structured knowledge base, knowledge entries corresponding to the first semantic slot set are retrieved to form a candidate set; The similarity between the first semantic slot set and the candidate set is calculated using the cosine similarity algorithm to obtain a similarity set; If any similarity in the similarity set is greater than a preset similarity threshold, the target field value is determined based on the candidate set corresponding to the current similarity. The target response text is obtained by querying the definition text corresponding to the target field value in a pre-established structured knowledge base and performing part-of-speech tagging on the definition text.

6. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, The target response text is subjected to spectral envelope construction and resonant peak modulation to obtain a unidirectional cue audio carrying harmonic distribution, including: The target response text is subjected to acoustic conversion processing to obtain an initial speech stream; The amplitude spectrum of each frame of the initial speech stream is calculated based on the Fast Fourier Transform algorithm, and the peak points of each formant in the amplitude spectrum are connected in sequence to form the spectral envelope of the corresponding frame. Linear predictive coding analysis is performed on the spectral envelope, and the corresponding pole positions are obtained by solving the linear prediction coefficients; According to the modulation target, the frequency and bandwidth of the resonance peak are modulated by adjusting the position of the poles to obtain the modulated spectral parameters; The encoder processes the modulated spectral parameters and analyzes the preset fundamental frequency trajectory to generate an excitation signal. The excitation signal is restored to a time-domain waveform to generate a unidirectional prompt audio carrying harmonic distribution.

7. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, Calculate the overlap rate between the one-way cue audio and the real-time audio stream, and perform frequency band shifting on the one-way cue audio based on the masking value and phase difference to obtain the target cue audio, including: The one-way prompt audio and the real-time audio stream are converted in the time-frequency domain based on the fast Fourier transform algorithm to obtain the first spectrum matrix and the second spectrum matrix; Extract the mid-frequency band of the one-way cue audio and the bandwidth of the real-time audio stream; Harmonic group analysis is performed on the first spectrum matrix and the second spectrum matrix using the mid-frequency band and the bandwidth to determine the overlap rate between the one-way prompt audio and the real-time audio stream; If the overlap rate is greater than the preset overlap value, then the auditory masking effect is calculated on the first spectrum matrix and the second spectrum matrix to obtain the masking value and phase difference; Based on the masking value and the phase difference, the mid-frequency band is shifted to obtain a target prompt audio with limited shift amplitude.

8. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 1, characterized in that, The target cues audio's spectral envelope and the harmonic distribution of the real-time audio stream are time-axis aligned and shifted with amplitude compensation using a mixing and overlay method to obtain and output the accompanying response audio, including: The target prompt audio and the real-time audio stream are processed using the Fast Fourier Transform algorithm to obtain the first amplitude spectrum and the second amplitude spectrum; The spectral envelope of the first amplitude spectrum and the harmonic distribution of the second amplitude spectrum are extracted to obtain the initial acoustic features; Based on the initial acoustic characteristics, the alignment offset of the spectral envelope and the harmonic distribution on the time axis is calculated using a dynamic time warping algorithm; The alignment offset is compared with a preset offset threshold, and combined with translation amplitude compensation, to obtain the compensated acoustic features. The compensated acoustic features are mixed and superimposed using a mixing track to obtain the accompanying response audio. The accompanying response audio is injected into the voice channel output based on the audio transmission protocol.

9. The intelligent voice interaction method based on semantic recognition and call linkage according to claim 8, characterized in that, Injecting the accompanying response audio into the voice channel output based on the audio transmission protocol includes: The accompanying response audio is parsed to obtain an audio state sequence; The audio state sequence is classified to determine the audio category label; If the audio category label meets the preset category conditions, the accompanying response audio is resampled according to the audio transmission protocol to obtain the first accompanying response audio; The first accompanying response audio is encoded to obtain an audio protocol data packet; Obtain the current buffer pool capacity of the target voice channel. If the current buffer pool capacity is greater than a preset capacity threshold, inject the audio protocol data packet into the target voice channel and obtain the injection status code. If the injected status code matches the preset success status identifier, a confirmation command for the interactive closed loop is generated and output.

10. An intelligent voice interaction system based on semantic recognition and call linkage, characterized in that, include: The text sequence acquisition module is used to acquire the real-time audio stream of multi-party voice calls and perform temporal feature extraction and acoustic decoding. Combined with the context window to mark the sentence boundaries and the semantic slots of the historical rounds, the text sequence is obtained. The entity recognition module is used to perform context entity recognition and dependency extraction on the fuzzy word frequency and intent category in the text sequence based on the conditional random field, so as to obtain a semantic feature vector carrying the semantic slot and triggering condition. The analysis module is used to comprehensively analyze the confidence level corresponding to the intent category and the triggering conditions to determine the target information requirement corresponding to the semantic feature vector. The retrieval and matching module is used to retrieve field values ​​and definition entries that match the semantic slots from a pre-established structured knowledge base according to the target information requirements, so as to obtain the target response text. The frequency domain analysis module is used to construct the spectral envelope and modulate the resonant peak position of the target response text to obtain a one-way prompt audio carrying harmonic distribution; The calculation module is used to calculate the overlap rate between the one-way prompt audio and the real-time audio stream, and to perform frequency band shifting processing on the one-way prompt audio by combining the masking value and phase difference to obtain the target prompt audio; The mixing and overlay output module is used to calculate the spectral envelope of the target cue audio and the alignment offset of the harmonic distribution of the real-time audio stream on the time axis. After processing by mixing and overlay, the accompanying response audio is obtained and output.