Sound interaction intention recognition and intelligent decision-making method based on AI large model

Through noise reduction processing and feature fusion technology, the problem of insufficient alignment of multimodal features in complex acoustic environments is solved, and higher intention recognition accuracy and device operation safety are achieved.

CN120279910AInactive Publication Date: 2025-07-08张婧
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510645816.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, in complex acoustic environments, multimodal features have insufficient fusion alignment due to noise interference and semantic segmentation errors, which affects the accuracy of intention recognition.

Method used

By collecting voice signals for noise reduction processing and feature extraction, combining band energy optimization and TF-IDF weight screening, fusion features are used to fusion mechanisms, and enhanced fusion feature vectors are generated based on the AI big model, intention reasoning and parameter verification are performed, and executable instructions are finally generated.

Benefits of technology

The multimodal feature fusion alignment is improved, the accuracy of intention recognition and the safety of equipment operation are improved, and the impact of computational load and noise interference is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279910A_ABST
    Figure CN120279910A_ABST
Patent Text Reader

Abstract

The invention discloses a sound interaction intention recognition and intelligent decision-making method based on an AI large model, and relates to the technical field of intelligent voice interaction, and the method comprises the steps: collecting a voice signal through an acoustic sensor, carrying out the noise reduction processing and acoustic feature extraction, capturing a text instruction, and carrying out the semantic segmentation and text feature extraction, splicing the acoustic feature vector and the text feature vector to form a multi-modal data packet; retrieving a historical memory library based on the enhanced fusion feature vector to generate a memory context vector, identifying the category of a deliberate map through a two-stage intention reasoning model, analyzing operation parameters, and outputting a structured intention instruction; and performing parameter legality verification, equipment state verification and security risk assessment on the structured intention instruction, and packaging the structured intention instruction into an executable instruction set after correcting abnormal parameters. According to the method, through double screening of the frequency band energy ratio and the lexical item importance score, the acoustic-text feature scale difference is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent voice interaction, and particularly to a method for identifying audio interaction intentions and making intelligent decisions based on an AI large model. Background Art

[0002] In recent years, intelligent voice interaction systems have made remarkable progress in the fields of acoustic signal processing and natural language understanding. In terms of speech signal processing, the WebRTC noise suppression (NS) algorithm and its improved schemes have become the mainstream noise reduction technologies, effectively suppressing environmental noise through frequency domain filtering and extracting acoustic features that conform to the auditory characteristics of the human ear by combining with Mel filter banks. At the same time, pre-trained language models (such as RoBERTa) have shown high accuracy in text clause segmentation and semantic understanding tasks, and can identify semantic boundaries in complex punctuation scenarios through context encoding. In the field of multi-modal fusion, attention mechanisms and dynamic weight allocation techniques are widely used for acoustic-text feature alignment, and the approximate nearest neighbor search (ANN) method based on Gaussian projection has improved the real-time performance in historical memory retrieval.

[0003] However, the existing technologies still face the following challenges: due to noise interference in acoustic signals, text semantic segmentation errors, and insufficient timeliness of historical memory, the cross-modal features result in low alignment of multi-modal feature fusion, thereby reducing the accuracy of intention recognition. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method for identifying audio interaction intentions and making intelligent decisions based on an AI large model to solve the problem of insufficient alignment of multi-modal features caused by noise interference and semantic segmentation deviation in a complex acoustic environment.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for identifying audio interaction intentions and making intelligent decisions based on an AI large model, which includes

[0008] collecting a voice signal through an acoustic sensor and performing noise reduction processing and acoustic feature extraction, and simultaneously capturing a text instruction for semantic clause segmentation and text feature extraction, and splicing the acoustic feature vector and the text feature vector to form a multi-modal data packet;

[0009] respectively performing frequency band energy optimization and TF-IDF weight screening on the acoustic feature vector and the text feature vector, fusing the features through an attention mechanism and dynamically adjusting the weights by superimposing the device state to generate a strengthened fusion feature vector;

[0010] Retrieve the historical memory bank based on the enhanced fusion feature vector to generate the memory context vector, identify the main intent category and parse the operation parameters through a two-stage intent inference model, and output the structured intent instruction;

[0011] Conduct parameter legality verification, device status verification, and security risk assessment on the structured intent instruction, and encapsulate it into an executable instruction set after correcting the abnormal parameters.

[0012] As a preferred solution of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model of the present invention, wherein: the noise reduction process includes the following steps,

[0013] Collect the analog voice signal and perform frame division processing to generate the spectrum matrix;

[0014] Analyze the spectrum matrix based on the improved WebRTC NS algorithm and calculate the real-time signal-to-noise ratio estimation value;

[0015] Map the real-time signal-to-noise ratio estimation value according to the dynamic threshold rule, perform frequency-domain filtering on the spectrum matrix, and reconstruct the denoised audio frame through inverse Fourier transform to generate the pure acoustic waveform data stream.

[0016] As a preferred solution of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model of the present invention, wherein: the acoustic feature extraction includes the following steps,

[0017] Frame the pure acoustic waveform data stream and perform fast Fourier transform, calculate the power spectrum energy distribution, and perform non-linear frequency scale conversion to generate the logarithmic Mel spectrum;

[0018] Perform time dimension averaging and normalization on consecutive frames of logarithmic Mel spectrum to form the acoustic feature vector.

[0019] As a preferred solution of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model of the present invention, wherein: the text feature extraction includes the following steps,

[0020] Preprocess the original text string and input it into the RoBERTa sentence splitting model to generate the sentence probability distribution sequence and identify the candidate semantic boundaries to generate the semantic unit sequence;

[0021] Extract the RoBERTa word embedding vector of the first character of the semantic unit and generate the normalized text feature vector through layer normalization.

[0022] As a preferred solution of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model of the present invention, wherein: the fusion of features through the attention mechanism includes the following steps,

[0023] Generate an acoustic feature preference mask matrix based on the band energy ratio of acoustic feature vectors to screen key band features;

[0024] Generate a binary mask matrix based on the TF-IDF values of text feature vectors to screen core semantic features;

[0025] Use the key band features as query vectors and the core semantic features as key value vectors, calculate and optimize the attention scores to generate an optimized attention map;

[0026] Read the device CPU load and network latency data in real time to generate an adjustment factor, dynamically correct the initial fusion weights and fuse them with the optimized attention map, and generate a strengthened fusion feature vector through global average pooling.

[0027] As a preferred solution of the method for identifying acoustic interaction intentions and making intelligent decisions based on the AI large model described in the present invention, wherein: the generation of the memory context vector includes the following steps,

[0028] Map the strengthened fusion feature vector to a hash code based on random projection hashing and retrieve similar records;

[0029] Calculate the cosine similarity for candidate records, generate a comprehensive weight in combination with the time decay weight, and perform weighted summation with the enhanced fusion feature vector to output the memory context vector.

[0030] As a preferred solution of the method for identifying acoustic interaction intentions and making intelligent decisions based on the AI large model described in the present invention, wherein: the two-stage intention reasoning includes the following steps,

[0031] Concatenate the memory context vector and the strengthened fusion feature vector to output the probability distribution of the main intention category;

[0032] Extract the band energy change rate based on the main intention in the category probability distribution and calculate the amplitude percentage;

[0033] Extract the device ID of the most recent operation from the historical records, integrate the direction, amplitude, and device parameters to generate a structured intention instruction.

[0034] As a preferred solution of the method for identifying acoustic interaction intentions and making intelligent decisions based on the AI large model described in the present invention, wherein: the encapsulation into an executable instruction set includes the following steps,

[0035] Iteratively predict the voltage increment based on the power amplifier parameters of the target device in the structured intention instruction;

[0036] When the predicted voltage exceeds the distortion threshold, trigger an alarm, and at the same time, count the user's historical rejection rate, and encapsulate the instructions that pass the verification into an executable instruction set.

[0037] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the method for identifying audio interaction intentions and making intelligent decisions based on an AI large model as described in the first aspect of the present invention is implemented.

[0038] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the method for identifying audio interaction intentions and making intelligent decisions based on an AI large model as described in the first aspect of the present invention is implemented.

[0039] The beneficial effects of the present invention are as follows: By adopting the improved WebRTC NS algorithm, the DSP computing load is reduced through non-uniform sub-band merging and automatic switching of the noise reduction mode based on the dynamic signal-to-noise ratio threshold; the long text sentence splitting accuracy is improved by the sentence splitting model based on RoBERTa-base through character-level probability distribution prediction and hierarchical backtracking strategy; the difference in the acoustic-text feature scale is reduced by the text semantic preference mask based on the relaxed acoustic band preference mask and TF-IDF cumulative weight truncation through double screening of the band energy ratio and term importance scoring; the fusion weight offset error is reduced by combining the dynamic adjustment of the attention weight based on device state perception. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0041] Figure 1 It is a flowchart of the method for identifying audio interaction intentions and making intelligent decisions based on an AI large model.

[0042] Figure 2 It is a schematic diagram of the improved WebRTC NS noise reduction processing.

[0043] Figure 3 It is a schematic diagram of acoustic feature extraction and text feature extraction.

[0044] Figure 4 It is a schematic diagram of enhanced feature fusion and intention reasoning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0046] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Persons skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0047] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.

[0048] Referring to Figures 1 to 4 , which is an embodiment of the present invention. This embodiment provides an audio interaction intention recognition and intelligent decision-making method based on an AI large model, including the following steps:

[0049] S1. Start the signal acquisition function of the acoustic sensor, configure the analog-to-digital converter to digitally sample the analog voice signal input by the microphone at a sampling rate of 16 kHz, and generate a pulse code modulation data stream containing continuous timestamps; perform time-domain normalization processing on the pulse code modulation data stream to generate a normalized time-domain waveform data stream;

[0050] Start the input event listening thread of the interaction interface to capture the text instruction byte stream input by the user through the physical keyboard or virtual keyboard in real time; perform UTF-8 encoding and decoding operations on the text instruction byte stream to convert the original byte sequence into a Unicode character sequence, and generate an unprocessed original text string;

[0051] Perform frame segmentation on the normalized time-domain waveform data stream, use the Hamming window function to perform sliding window segmentation with a fixed frame length and frame shift, and generate a set of audio frames containing a time series; perform a fast Fourier transform on each audio frame in the set of audio frames to convert the time-domain signal into a frequency-domain energy spectrum, and generate a spectrum matrix containing the amplitude of frequency components.

[0052] Further explanation: According to the Nyquist sampling theorem, the sampling rate needs to be greater than 2 times the highest frequency of the signal. An 8 kHz sampling rate can only cover the 0 - 4 kHz frequency band and cannot capture Chinese voiceless fricatives. A 16 kHz sampling rate can cover 0 - 8 kHz and completely retain the core frequency band of Mandarin Chinese. The 16 kHz sampling rate and the hybrid normalization strategy achieve the optimal balance in Chinese speech recognition tasks.

[0053] It should be noted that covering the sensitive frequency band of the human ear (Nyquist frequency of 20Hz - 8kHz) can balance the computational load and voice fidelity; through Hamming window framing, spectral leakage can be reduced and frequency domain resolution can be improved; time domain conversion can eliminate the microphone gain difference, standardize the signal amplitude, and avoid numerical overflow in subsequent processing.

[0054] Analyze the spectral matrix based on the improved WebRTC NS algorithm, calculate the ratio of the current ambient noise floor energy to the voice signal energy, and generate a real-time signal-to-noise ratio estimate value.

[0055] According to the dynamic threshold adjustment rule, map the real-time signal-to-noise ratio estimate value within the dynamic threshold range. For example:

[0056] When the signal-to-noise ratio < -8dB, activate the deep noise reduction mode.

[0057] When -8dB ≤ signal-to-noise ratio ≤ +3dB, use linear interpolation to calculate the dynamic threshold.

[0058] When the signal-to-noise ratio > +3dB, turn off the active noise reduction.

[0059] Apply the dynamic threshold to perform frequency domain filtering on the spectral matrix, retain the frequency components higher than the current threshold, and generate the frequency domain energy spectrum after noise reduction; perform inverse fast Fourier transform on the frequency domain energy spectrum after noise reduction to restore the frequency domain signal to a time domain waveform and generate a denoised audio frame; use the overlap-add method to reconstruct the time domain signal of all denoised audio frames and generate a pure acoustic waveform data stream with continuous timestamps.

[0060] Furthermore, the critical values of -8dB and +3dB are set based on the result of the balance between the human ear auditory masking effect and speech clarity. When the ambient noise energy exceeds the voice signal by 8dB, the human ear's recognition of speech will drop to the critical point, and at this time, it is necessary to start a 16th-order Butterworth filter bank for full-band noise suppression; when the signal-to-noise ratio exceeds 3dB, the improvement in sound quality brought by active noise reduction is less than 0.3 MOS points, but the power consumption increases. At this time, turning off the noise reduction can save DSP computing power.

[0061] It should be noted that WebRTC NS defaults to using 32 sub-band frequency domain processing, and each frame requires 2048-point FFT calculation, resulting in a high DSP load rate. While the improved WebRTC NS combines sub-bands into non-uniform sub-bands, reduces the FFT points, decreases the computational amount, converts the floating-point operation to Q15 fixed-point format, reduces the occupancy of multiplier resources, and accelerates matrix operations using the SIMD instruction set (such as ARM NEON), reducing the real-time delay.

[0062] Perform frame segmentation on the denoised pure acoustic waveform data stream, and use a rectangular window with a fixed duration to perform non-overlapping segmentation and cutting to generate a discrete audio frame sequence; perform a fast Fourier transform on each audio frame in the discrete audio frame sequence to calculate the power spectrum energy distribution in the linear frequency scale;

[0063] In the human ear sensitive frequency band from 20 Hz to 4000 Hz, deploy a Mel filter bank composed of triangular band-pass filters to perform non-linear frequency scale conversion on the power spectrum energy distribution and output the energy value; perform logarithmic compression processing on the output energy value of each filter to generate a logarithmic Mel spectrum that conforms to the human ear auditory characteristics;

[0064] Perform time dimension average calculation on the logarithmic Mel spectra of multiple consecutive audio frames to generate a Mel spectrogram mean matrix with temporal smoothing characteristics; perform L2 norm normalization on the Mel spectrogram mean matrix along the frequency axis to eliminate the energy dimension difference, and extract the numerical values of the frequency channels of the normalized Mel spectrogram mean matrix and arrange them in order to form an acoustic feature vector.

[0065] It should be noted that by deploying the Mel filter bank, the noise reduction intensity can be dynamically adjusted, burst noise suppression can be performed, and double breakthroughs in noise reduction performance and calculation efficiency are achieved.

[0066] Perform preprocessing operations on the original text string, remove leading and trailing spaces and consecutive repeated spaces, convert full-width characters to half-width forms to generate a normalized text string; input the normalized text string into a pre-trained RoBERTa sentence splitting model, and output the sentence splitting probability values at each character position through the classification head at the end of the model to generate a sentence splitting probability distribution sequence;

[0067] Identify the peak points with probability values exceeding the sentence splitting probability threshold in the sentence splitting probability distribution sequence to determine the candidate semantic boundary positions and generate a sentence splitting position index list.

[0068] Further explanation, the sentence splitting probability threshold is set based on the PR curve analysis of the RoBERTa-base in the Chinese sentence splitting task validation set.

[0069] Perform cutting operations on the normalized text string according to the sentence splitting position index list, and insert segmentation marks at each sentence splitting position to generate an initial semantic unit fragment set;

[0070] Perform length verification on the initial semantic unit fragment set, for example:

[0071] If the number of fragment characters ≤ 128, directly retain it as the final semantic unit;

[0072] If the number of fragment characters > 128, trace back to the nearest comma, period or exclamation mark position for secondary segmentation.

[0073] Arrange all the semantic units that pass the verification in the order of the original text to generate a sequence of semantic units that meet the length constraints; traverse each semantic unit in the sequence of semantic units to locate the byte offset position of the first character in the original normalized text string; input the Unicode code point of the first character into the word embedding lookup table of the pre-trained RoBERTa sentence splitting model, retrieve the corresponding character embedding vector, and perform layer normalization on the character embedding vector to eliminate the scale difference in the embedding space and generate a normalized text feature vector; perform a dimensional concatenation operation on the acoustic feature vector and the normalized text feature vector, and connect them along the feature axis to form a joint feature vector.

[0074] Furthermore, based on the pre-trained language model of RoBERTa-base, through the context encoding of the hidden layer, it can accurately distinguish the corresponding scenarios, such as full stop ambiguity and exclamation mark emotion enhancement, etc.; through

[0075] The triple technological innovation of RoBERTa context modeling, dynamic threshold segmentation and hierarchical hybrid strategy improves the long text processing efficiency on the premise of ensuring the sentence splitting accuracy; the character embedding normalization technology improves the cross-modal feature alignment degree and lays a foundation for subsequent multi-modal fusion.

[0076] It should be noted that the hybrid segmentation strategy includes three layers of segmentation. The first segmentation uses RoBERTa to output the sentence splitting probability of each position and mark the candidate segmentation points; the second segmentation searches backward from the end for the nearest secondary punctuation marks such as commas and semicolons for segments with a length > 128 characters. If the backtracking fails, the length is allowed to be extended to 256 characters; the third segmentation forces short sentences through dependency syntax, uses the LTP toolkit to analyze the syntax tree, and forces segmentation at the core predicate or object boundary; there is a scale difference between the character embedding of RoBERTa and other modal features (such as acoustic vectors), and it is matched with the acoustic features by calculating the statistical mean / standard deviation.

[0077] Read the joint feature vector and verify that the dimension integrity meets the input requirements of the trainable parameter matrix; perform a linear transformation on the joint feature vector and the trainable parameter matrix to generate an intermediate feature vector; perform an element-wise addition operation on the intermediate feature vector by superimposing the bias term to generate a bias-corrected feature vector; apply the sigmoid activation function to perform a non-linear mapping on the bias-corrected feature vector to generate an initial modal fusion weight vector; pack the acoustic feature vector, the sequence of semantic units and the initial modal fusion weight vector into a multi-modal data packet.

[0078] It should be noted that by concatenating acoustic features (such as Mel spectrum energy distribution) and text features (such as RoBERTa character embeddings) along the feature dimension, the model can capture the correlation between prosodic variations in speech (such as the rising intonation at the end of an interrogative sentence) and semantic markers in the text (such as exclamation marks and question marks). For example, when the fundamental frequency of the speech significantly rises at the end of a sentence, the question mark in the text will be enhanced and weighted, thereby improving the recognition accuracy of the interrogative intent; the Sigmoid function is used to limit the fusion weight between 0 and 1 to prevent a single modality from overly dominating the decision-making. For example, when there are spelling mistakes or garbled characters in the text, the acoustic weight is automatically increased to avoid interference from incorrect text.

[0079] S2. Analyze the Mel frequency scale corresponding to the acoustic feature vector, and based on the center frequency parameters of the Mel filter bank, establish a mapping relationship table between the Mel band numbers and the linear frequencies; screen out all Mel band numbers whose center frequencies are in the range of 20 Hz to 4000 Hz in the mapping relationship table to generate a list of valid band indices; extract the element values corresponding to the valid band indices in the acoustic feature vector to obtain a set of energy values in the 20 - 4000 Hz frequency band; perform a summation operation on the set of energy values to calculate the total energy value, and perform normalization processing on the energy value of each valid band to calculate the energy proportion, and the expression is:

[0080]

[0081] where, p i is the energy proportion of the i-th band, used to quantify the relative importance of this band, e i is the original energy value of the i-th band, reflecting the absolute energy intensity of the band in the acoustic signal, A is the total energy sum of all valid bands, providing a normalization benchmark to ensure the comparability of the calculation of the energy proportion, ∈ is a minimum value, a numerical stability term to prevent the denominator from being zero and avoid division-by-zero errors.

[0082] Arrange all the energy proportion values according to the order of the valid band indices to generate a band energy distribution vector aligned with the dimension of the original acoustic feature vector.

[0083] Perform a descending order operation on the band energy distribution vector to generate an index sequence sorted from high to low according to the energy proportion, calculate the cumulative energy proportion value of the sorted elements, and construct an initial binary mask vector; apply the Gumbel-Softmax relaxation technique to perform a differentiable approximation on the initial binary mask vector to generate a relaxed mask vector, and the expression is:

[0084]

[0085] where, m iis the differentiable mask value (continuous probability) for the i-th frequency band, which is a continuous relaxation form of the binary mask, retaining gradients to support backpropagation, g i is the noise for the i-th frequency band sampled from the standard Gumbel distribution (location parameter 0, scale parameter 1), making the mask selection process differentiable and approximating the discrete sampling behavior. τ is the temperature coefficient that controls the sharpness of the probability distribution: when τ approaches 0, it approximates discrete selection; when τ approaches ∞, it tends to a uniform distribution, p j is the energy proportion of the j-th frequency band, g j is the noise for the j-th frequency band sampled from the standard Gumbel distribution;

[0086] Perform truncation on the relaxed mask vector: set the elements with differentiable mask values ≥ the symmetric splitting point of the binary decision to 1, and the rest to 0, generating the final preferred mask matrix for acoustic features.

[0087] Furthermore, the Mel filter bank is designed based on the non-linear characteristics of human ear hearing, and the center frequencies are distributed in logarithmic intervals. In the range of 20Hz to 4000Hz, dense filters are set in the low-frequency region (such as 20 - 1000Hz) to capture the details of the fundamental frequency and formants, and the high-frequency region (1000 - 4000Hz) is gradually thinned out to simulate the decline of the human ear's high-frequency resolution. An exponential decay strategy is used to adjust the temperature coefficient: initially, τ = 1.0, making the mask probability close to a uniform distribution to promote the model to explore the importance of different frequency bands; as the number of training rounds increases, the temperature coefficient is gradually reduced to strengthen the certainty of high-frequency band selection; according to the real-time calculated attention score distribution, the 10th percentile is dynamically selected as the threshold benchmark. For example, when the environmental signal-to-noise ratio (SNR) is below 0dB, the percentile threshold automatically drops to the 5th percentile to retain more weak speech components; when the SNR is above 10dB, it returns to the 15th percentile to filter out redundant noise. In a burst noise scenario (such as keyboard tapping), the speech intelligibility (STOI) is improved.

[0088] Arrange all the semantic units in the semantic unit sequence in the order of appearance, construct a document set containing all the semantic unit texts, and perform word segmentation on each semantic unit, count the number of times each term appears in the corresponding semantic unit, and calculate the term frequency. The expression is:

[0089]

[0090] where TF(t, d) is the frequency of term t in document d;

[0091] Statistical distribution of each term in the entire document set, calculate the inverse document frequency, the expression is:

[0092]

[0093] Among them, IDF(t) is the frequency of term t in the inverse document.

[0094] Perform weight calculation on the terms in each semantic unit to obtain the importance score of term t in document d. The expression is:

[0095] TF-IDF(t, d) = TF(t, d) × IDF(t);

[0096] Among them, TF-IDF(t, d) is the importance score of term t in document d.

[0097] Perform weighted summation on the TF-IDF values of all terms within each semantic unit to calculate the overall weight of the semantic unit. The expression is:

[0098] w d = ∑ t∈d TF-IDF(t, d);

[0099] Among them, w d is the overall weight of the semantic unit.

[0100] Sort the weight values of all semantic units in descending order, calculate the critical position when the cumulative weight sum reaches 30% of the total weight, determine the number of retained semantic units, construct a binary vector with the same dimension as the semantic unit sequence, assign 1 to the corresponding positions of the retained semantic units, and assign 0 to the rest to generate a text feature preference mask.

[0101] Further explanation: When calculating the inverse document frequency, the Laplace smoothing technique is adopted, adding a constant to the denominator to avoid division-by-zero errors caused by out-of-vocabulary words (terms that do not appear in any document), and improving the numerical stability of the algorithm in small-scale datasets and sparse data scenarios; for different vertical domains, a hybrid word segmentation strategy is adopted: in general scenarios (such as social media), a pre-trained word segmentation model based on BERT is used to capture context-sensitive word segmentation boundaries; in professional fields, a domain term library and a rule engine are combined to ensure the complete retention of professional nouns. On the basis of traditional TF-IDF, a part-of-speech weight coefficient is introduced, and through weighted cumulative weight truncation, key information is accurately extracted in the dialogue, reducing the missed detection rate of diagnostic key clues.

[0102] Perform element-wise multiplication on each element of the acoustic feature preference mask matrix and the corresponding element of the original acoustic feature vector to extract key frequency band features; perform element-wise multiplication on each element of the text feature preference mask matrix and the corresponding element of the standardized text feature vector to extract core semantic features.

[0103] Use the key frequency band features as the query vector and the core semantic features as the key-value vector to calculate the attention score. The expression is:

[0104]

[0105] Among them, A is the attention score, Q is the query vector, K is the key vector, T is the transpose, and v is the scaling factor, that is, the dimension of the key vector.

[0106] To further illustrate, by expanding the attention scores (e.g., reorganizing the timing, frequency band, and semantic channel dimensions), the cross-channel-space features are simultaneously aggregated in the fusion stage. For example, in in-vehicle voice interaction, the association between the energy change of the acoustic frequency band (spatial dimension) and the timing of the occurrence of text sentiment words (channel dimension) is captured simultaneously, and the global average pooling further compresses the number of parameters, retaining only the mean features of each dimension.

[0107] In the attention score matrix, the attention score threshold is set based on the quantile of the attention score distribution, and the elements below the attention score threshold are set to zero to generate an optimized attention map; at the same time, the CPU load rate and network delay data of the device are read in real time and combined into a two-dimensional state vector; the two-dimensional state vector is input into a multilayer perceptron containing two linear layers, and the adjustment factor is output through two matrix transformations and two LeakyReLU activation functions; the initial modal fusion weight is element-wise multiplied with the adjustment factor to generate the final fusion weight that reflects the device state; the optimized attention map is expanded into a three-dimensional tensor, and cross-dimensional fusion calculation is performed with the final fusion weight tensor, and dimensionality reduction is performed through global average pooling to form an enhanced fusion feature vector.

[0108] S3. Extract all interaction records stored in the last 24 hours from the multimodal memory library to generate a list of interaction records with timestamps; create a random projection matrix based on the feature vector dimension of the interaction record list, where each element in the random projection matrix is ​​independently sampled from a standard Gaussian distribution (mean 0, variance 1) to generate a projection parameter matrix; input the enhanced fusion feature vector as the query vector into the projection parameter matrix, perform the same linear projection and independent sampling, and generate a query hash code; extract the first 8 binary bits of the query hash code and convert it into a decimal bucket number; retrieve all record entries under the corresponding bucket number from the hash bucket index table, remove records with timestamps greater than 24 hours, and generate a preliminary candidate record set; extract the enhanced fusion feature vector of each record from the preliminary candidate record set to calculate the cosine similarity with the query vector; sort the candidate records from high to low according to the cosine similarity value to generate a sorted similarity list; intercept the first N similar memory records in the similarity list to generate a final similar memory set.

[0109] Further explanation: The setting of the 24-hour time window is based on the statistical law of user behavior analysis: 85% of smart device interaction behaviors have strong correlation within 24 hours (such as playing the same playlist continuously and adjusting the volume repeatedly). By automatically removing overdue records and eliminating the interference of outdated preferences (such as the night mode set yesterday) on current decisions, the accuracy of memory retrieval is improved (measured smart home data set). The dimension of the Gaussian random projection matrix is ​​the same as the original feature vector. By retaining the mathematical properties of the inner product relationship, the similarity between the vectors after dimensionality reduction is ensured to be consistent with the original space.

[0110] For the first N similar memory records in the final similar memory set, extract the storage timestamps respectively, and calculate the difference between the current time and the timestamp of each record; apply the exponential decay formula to the time difference of each record to generate the time decay weight, which is expressed as:

[0111]

[0112] Among them, α u is the timeliness weight of the u-th memory record. The larger the value, the newer the record and the higher the weight. u is the time difference of the u-th memory record, which quantifies the newness of the record. The larger the time difference, the older the record.

[0113] The similarity value is multiplied by the time decay weight to generate a comprehensive weight, and the temperature coefficient is applied to the comprehensive weight for scaling to obtain a normalized weight. The enhanced fusion feature vectors of the first N similar memory records are extracted, and the normalized weight is multiplied element by element with the corresponding feature vector in the enhanced fusion feature vector, and then the sum is obtained to obtain a memory context vector.

[0114] Further explanation: The 24-hour half-life is based on the Ebbinghaus forgetting curve, which simulates the natural decay of human memory. The minimum weight threshold for retaining long-tail records after normalization ensures that unpopular operations (such as "turn on the air purifier") can still be triggered in specific scenarios (such as when PM2.5 exceeds the standard).

[0115] The memory context vector is subjected to two-stage intent reasoning, where the first stage performs coarse-grained intent classification and the second stage performs fine-grained parameter parsing, as follows:

[0116] The first stage: concatenate the enhanced fusion feature vector and the memory context vector along the feature axis to generate an enhanced feature vector; input the enhanced feature vector into the classification head of the MiniLLM-3B model after knowledge distillation optimization to generate the original score vector of the preset intent category.

[0117] It should be noted that the preset intent categories include volume adjustment, playback control, mode switching, etc.

[0118] Perform Softmax normalization on the original score vector to generate the probability distribution of intent categories, arrange them in descending order, and select the category label corresponding to the maximum probability value after sorting as the main intent.

[0119] Further explanation, the high-frequency band is clearly defined as 4000 - 8000Hz, covering the sensitive area of the human ear to speech intensity and environmental noise. This frequency band is selected because the change of speech energy above 4000Hz can better reflect the user's vocal intensity.

[0120] The second stage: Based on the main intent category label (such as volume adjustment), activate the corresponding parameter parsing logic branch; extract the energy value of the frequency band above 4000Hz from the mean matrix of the Mel spectrogram to generate a high-frequency energy time series, and calculate the mean change rate of the high-frequency energy of the nearest time frame. The expression is:

[0121]

[0122] Among them, ΔE is the mean change rate of high-frequency energy, E cur is the mean high-frequency energy of the current time frame, E pre is the mean high-frequency energy of the previous time frame;

[0123] If ΔE≥0.1, increase the judgment direction; if ΔE≤-0.1, decrease the judgment direction; if -0.1<ΔE<0.1, inherit the default direction of the main intent in the first stage;

[0124] Extract the L2 norm of the current Mel spectrogram mean matrix and the L2 norm of the previous moment, calculate the relative change rate, and convert it into an amplitude percentage. The expression is:

[0125]

[0126] Among them, Y is the amplitude percentage, L cur is the current L2 norm, L pre is the L2 norm of the previous moment;

[0127] Based on the amplitude percentage, constrain the amplitude range. For example: truncate the result to 1% - 100%, with a step size of 5% (such as the calculated value 37% → 35%, 42% → 40%).

[0128] Further explanation, the L2 norm calculates the overall energy of the Mel spectrogram matrix, reflecting the global intensity change of the speech signal, rather than a single frequency peak, avoiding misjudgment of peaks caused by sudden noises (such as brief coughs). Compared with the spectral peak method, the L2 norm has improved accuracy in identifying the volume direction in a noisy environment.

[0129] Extract the list of speaker device IDs that have been operated within the last 24 hours from the interaction records associated with the memory context vectors, sort them in descending order of the operation timestamp, and select the device ID with the most recent time as the target device. If there is no historical record, by default, select the device currently connected to the interaction interface as the target device;

[0130] Integrate the main intent type, operation direction, amplitude range, and target device to generate a structured intent instruction.

[0131] S4. Extract the amplitude percentage field from the structured intent instruction (e.g., the amplitude is 15%), and convert it into a numerical form; according to the target device ID, read the maximum gain limit value allowed by the device from the device specification library:

[0132] If the amplitude percentage > the maximum gain limit: Mark the parameter as illegal, trigger the correction logic, force the amplitude to be corrected to the maximum gain limit value (e.g., 50%), and update the amplitude field in the instruction;

[0133] If the amplitude percentage ≤ the maximum gain limit: Retain the original parameter and mark it as legal.

[0134] Obtain the list of IDs of all online devices from the current network topology (e.g., the main speaker in the living room, the bedroom audio).

[0135] Check whether the target device ID in the intent instruction exists in the online device list:

[0136] If it exists: Retain the original device ID and mark it as legal;

[0137] If it does not exist: Retrieve the device records that have been successfully operated within the last 24 hours from the multimodal memory library, sort them in descending order of the operation timestamp, extract the device ID in the first record as a replacement. If there is no valid record in the memory library, extract the default device ID currently connected to the interaction interface (e.g., the main speaker in the living room) and update the device field in the intent instruction.

[0138] Extract the direction field from the structured intent instruction (e.g., the direction is +), and convert it into a direction symbol identifier (e.g., +); verify whether the direction identifier belongs to the direction set (e.g., +, -):

[0139] If it conforms: Retain the direction parameter and mark it as legal;

[0140] If it does not conform (e.g., unknown): Mark the parameter as illegal, trigger the exception handling, and inherit the default direction of the main intent in the first stage (e.g., the default direction of "volume adjustment" is "+").

[0141] Read the power amplifier circuit parameters of the target device from the device specification library, including the power amplifier time constant, the maximum safe output voltage, and the distortion threshold; read the current output voltage in real time, calculate the target voltage increment in combination with the voltage amplitude percentage of the intent instruction, and iteratively predict the voltage value according to the discrete time step; when the voltage change rate of the time steps of three consecutive predicted voltage values no longer fluctuates, determine that the stable state is entered and record it as the predicted stable voltage value; calculate the difference between the predicted stable voltage and the distortion threshold. If the difference > 0, mark that there is a distortion risk and trigger correction or alarm; if the difference ≤ 0, mark it as safe and allow execution.

[0142] Extract the target device ID in the current intent instruction, filter the records with exactly the same operating device ID from the user historical preference library, and filter the historical operation records that are exactly the same as the current intent type (such as "volume adjustment"); extract the user response field (execute or reject) of each record, count the total execution times (the number of all records) and the rejection times (the number of records with the user response being reject), and calculate the rejection rate (rejection times / total execution times); compare the calculated rejection rate with the preset threshold of 60%:

[0143] If the rejection rate ≥ 60%: Trigger the secondary confirmation process and generate a natural language prompt containing the reason for the conflict (such as "60% of similar operations in history were rejected. Do you want to continue?")

[0144] If the rejection rate < 60%: Skip the confirmation, mark the current instruction as passed verification, and enter the execution queue.

[0145] Encapsulate all the instructions that pass the verification in JSON format to generate an executable instruction set.

[0146] This embodiment also provides a computer device applicable to the case of the sound interaction intent recognition and intelligent decision-making method based on the AI large model, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the sound interaction intent recognition and intelligent decision-making method proposed in the above embodiment.

[0147] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0148] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing audio interaction intention recognition and intelligent decision-making based on an AI large model as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, magnetic disks, or optical discs.

[0149] In summary, the present invention reduces the DSP computing load by adopting an improved WebRTC NS algorithm, automatically switching the noise reduction mode through non-uniform sub-band merging and dynamic signal-to-noise ratio thresholds; improves the long text sentence splitting accuracy through a sentence splitting model based on RoBERTa-base with character-level probability distribution prediction and hierarchical backtracking strategy; reduces the acoustic-text feature scale difference through a relaxed acoustic frequency band preference mask and a text semantic preference mask with TF-IDF cumulative weight truncation, through dual screening of the frequency band energy ratio and the term importance score; and reduces the fusion weight offset error by combining dynamic adjustment of attention weights based on device state perception.

[0150] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An audio interaction intention recognition and intelligent decision-making method based on a large AI model, characterized in that: including collecting a voice signal through an acoustic sensor, performing noise reduction processing and acoustic feature extraction, simultaneously capturing a text instruction for semantic clause segmentation and text feature extraction, and splicing an acoustic feature vector and a text feature vector to form a multimodal data packet; respectively performing band energy optimization and TF-IDF weight screening on the acoustic feature vector and the text feature vector, fusing features through an attention mechanism and dynamically adjusting weights by superimposing device states to generate a reinforced fusion feature vector; retrieving a historical memory bank based on the reinforced fusion feature vector to generate a memory context vector, identifying a main intention category and parsing operation parameters through a two-stage intention inference model, and outputting a structured intention instruction; performing parameter legality verification, device state verification and security risk assessment on the structured intention instruction, and encapsulating it into an executable instruction set after correcting abnormal parameters.

2. The method for identifying audio interaction intentions and making intelligent decisions based on the AI large model according to claim 1, characterized in that: The noise reduction processing includes the following steps collecting an analog voice signal and performing frame processing to generate a spectrum matrix; analyzing the spectrum matrix based on an improved WebRTC NS algorithm to calculate a real-time signal-to-noise ratio estimation value; mapping the real-time signal-to-noise ratio estimation value according to a dynamic threshold rule, performing frequency-domain filtering on the spectrum matrix, and reconstructing a denoised audio frame through inverse Fourier transform to generate a pure acoustic waveform data stream.

3. The method for identifying sound interaction intentions and making intelligent decisions based on the AI large model according to claim 1, wherein: The acoustic feature extraction includes the following steps framing the pure acoustic waveform data stream and performing fast Fourier transform, calculating the power spectrum energy distribution, and performing non-linear frequency scale conversion to generate a logarithmic Mel spectrum; performing time dimension averaging and normalization on consecutive multi-frame logarithmic Mel spectra to form an acoustic feature vector.

4. The method for identifying sound interaction intentions and making intelligent decisions based on the AI large model according to claim 1, wherein: The text feature extraction includes the following steps preprocessing an original text string and inputting it into a RoBERTa clause segmentation model to generate a clause probability distribution sequence and identify candidate semantic boundaries, generating a semantic unit sequence; extracting RoBERTa word embedding vectors of the first characters of semantic units, and generating a normalized text feature vector through layer normalization.

5. The method for identifying sound interaction intentions and making intelligent decisions based on the AI large model according to claim 1, wherein: The fusing features through the attention mechanism includes the following steps generating an acoustic feature optimization mask matrix based on the band energy ratio of the acoustic feature vector to screen key band features; generating a binary mask matrix based on the TF-IDF value of the text feature vector to screen core semantic features; using the key band features as query vectors and the core semantic features as key-value vectors, calculating attention scores and optimizing them to generate an optimized attention map; real-time reading device CPU load and network latency data to generate an adjustment factor, dynamically correcting the initial fusion weight and fusing it with the optimized attention map, and generating a reinforced fusion feature vector through global average pooling.

6. The method for identifying audio interaction intent and making intelligent decisions based on the AI large model according to claim 1, wherein: The generating the memory context vector includes the following steps mapping the reinforced fusion feature vector to a hash code based on random projection hashing and retrieving similar records; calculating the cosine similarity of candidate records, generating a comprehensive weight by combining time decay weights, and performing weighted summation with the enhanced fusion feature vector to output a memory context vector.

7. The method for identifying audio interaction intentions and making intelligent decisions based on the AI large model according to claim 1, characterized in that: The two-stage intention inference includes the following steps splicing the memory context vector and the reinforced fusion feature vector to output a main intention category probability distribution; extracting the band energy change rate based on the main intention in the category probability distribution and calculating the amplitude percentage; Extract the device ID of the most recent operation from the historical record, integrate the direction, amplitude, and device parameters to generate a structured intent instruction.

8. The method for recognizing audio interaction intentions and making intelligent decisions based on the AI large model according to claim 1, characterized in that: The encapsulation into an executable instruction set includes the following steps: Iteratively predict the voltage increment based on the target device power amplifier parameters in the structured intent instruction; When the predicted voltage exceeds the distortion threshold, trigger an alarm, simultaneously count the user's historical rejection rate, and encapsulate the instructions that pass the verification into an executable instruction set.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method for identifying audio interaction intent and making intelligent decisions based on the AI large model according to any one of claims 1 to 8.

Citation Information

Cited By

  • Voice interaction optimization method and system based on multi-modal large model

    CN120496511A

  • Intelligent device voice interaction method and system based on multi-task dialect recognition

    CN120766677A

  • Vehicle-mounted multi-modal agent construction method and device based on MCP

    CN120949926A

  • Routing method for multi-modal problem under AI platform, medium and system

    CN121009497A

  • Multi-round interactive AI Agent agent based on multi-module collaboration and implementation method thereof

    CN121009985A