Speech data processing method and system based on large language model
By employing a large language model-based speech data processing method, the speech understanding problem in complex acoustic environments within banking scenarios was solved, achieving high-precision semantic understanding and ambiguity resolution, and reducing business misjudgment rates and risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG JINWAN INFORMATION TECH CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-09
AI Technical Summary
Voice interaction in banking scenarios faces the challenge of accurately separating acoustic features in complex acoustic environments, resulting in low accuracy in understanding highly ambiguous speech and a lack of dynamic reasoning capabilities to address semantic uncertainty risks, thus increasing business misjudgments and risks.
We employ a speech data processing method based on a large language model. We construct an acoustic state representation tensor through multi-channel feature decomposition, generate a semantic candidate distribution space and construct a semantic evolution path graph, dynamically construct inference depth control parameters, execute a multi-path hypothesis parallel inference mechanism, construct an intent structure vector and map it into structured semantic output.
It improves the semantic understanding accuracy of voice interaction, reduces the business misjudgment rate and risk, and enhances the ability to resolve semantic ambiguity.
Smart Images

Figure CN122177097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech data processing technology, specifically to a speech data processing method and system based on a large language model. Background Technology
[0002] In core scenarios such as customer service, risk control, internal compliance, and business processing in banks, voice interaction is key to improving operational efficiency and user experience. Currently, the banking sector widely utilizes systems such as intelligent voice customer service, telephone banking transactions, credit review recording analysis, and compliance dual recording (audio and video recording) checks. However, voice interaction in banking scenarios is highly specialized, complex, and risk-sensitive, posing significant challenges to existing bank voice data processing. First, the banking voice interaction environment is complex. Users may initiate calls in noisy lobby areas, open-plan offices, or through unstable mobile networks, resulting in a large amount of environmental noise, channel interference, and echoes mixing into the original voice signal. This affects the accuracy of voice recognition, distorts semantic understanding, and leads to incorrect interpretation of the user's true intent, potentially causing serious business risks. Second, semantic expression in the banking sector is highly ambiguous and specialized. The same voice command may represent completely different financial operations in different contexts. For example, "Check my money" may refer to checking the current account balance, the total amount of time deposits, or even the market value of wealth management products. Current technology struggles to accurately interpret the prosodic structure of speech and the meaning of the words. Accurate capture of semantic ambiguity by speaker characteristics makes it impossible to construct accurate user intent, leading to frequent interaction errors or business interruptions. Furthermore, banking operations have extremely high requirements for security and compliance. When processing voice commands involving customer privacy, fund transactions, or investment intentions, it is necessary to strictly verify potential semantic ambiguity. Existing voice processing uses fixed inference paths and computing resource allocation, which cannot dynamically adjust the granularity of deep analysis according to the uncertainty of the voice signal. When faced with low-quality voice or high-risk intent, there is a lack of flexible multi-path verification mechanisms, which cannot effectively identify and eliminate erroneous assumptions caused by voice disturbances or misunderstandings, increasing the risk of erroneous transactions and violations.
[0003] Therefore, current technologies suffer from technical problems such as difficulty in accurately and effectively separating acoustic features in complex acoustic environments, resulting in low accuracy in understanding highly ambiguous speech, and a lack of dynamic reasoning capabilities to address the risk of semantic uncertainty. Summary of the Invention
[0004] This application provides a speech data processing method and system based on a large language model, which solves the technical problems in the prior art, such as the difficulty in accurately and effectively separating acoustic features in complex acoustic environments, resulting in low accuracy of understanding highly ambiguous speech, and the lack of dynamic reasoning ability to cope with semantic uncertainty risks. It achieves the technical effects of improving the semantic understanding accuracy and ambiguity resolution ability of speech interaction, and reducing the business misjudgment rate and risks.
[0005] This application provides a speech data processing method based on a large language model. The method includes: after receiving the original speech signal, performing multi-channel feature decomposition on the original speech signal to construct an acoustic state representation tensor containing acoustic perturbation components, prosodic structure components, and speaker voice embedding components; constructing a semantic candidate distribution space based on the acoustic state representation tensor, generating multiple sets of semantic hypothesis vectors for speech segments within different time windows, and constructing a semantic evolution path graph according to speech continuity constraints, wherein the semantic evolution path graph includes semantic nodes and temporal dependency weights between nodes; performing uncertainty calculation on the semantic evolution path graph to generate a semantic uncertainty function representing semantic ambiguity and speech perturbation sensitivity; dynamically constructing inference depth control parameters based on the semantic uncertainty function, and inputting the inference depth control parameters into a multi-layer inference path scheduling unit of the large language model to perform output competition selection under a multi-path hypothesis parallel inference mechanism to construct an intent structure vector; and mapping the intent structure vector to a structured semantic output.
[0006] In a possible implementation, constructing an intent structure vector includes: decomposing the semantic uncertainty function into semantic ambiguity gradient components and acoustic perturbation coupling components to construct an uncertainty layering matrix; calculating the inference complexity score within the corresponding time window based on the uncertainty layering matrix, and mapping the inference complexity score to a multi-level inference depth level according to a preset segmented threshold interval; generating differentiated inference resource allocation vectors for different inference depth levels, the inference resource allocation vectors including the attention head activation ratio, the number of inference layer expansions, and the number of parallel hypothesis paths; inputting the differentiated inference resource allocation vectors as inference depth control parameters into the multi-layer inference path scheduling unit of the large language model; constructing a multi-layer inference scheduling graph based on the multi-layer inference path scheduling unit, mapping different hypothesis paths to independent inference subgraph structures, and dynamically controlling the activation range and iteration depth of each inference subgraph according to the differentiated inference resource allocation vectors; during the multi-path hypothesis parallel inference process, constructing a cross-path consistency tensor, using the cross-path consistency tensor to jointly calculate the semantic consistency score, context continuity score, and perturbation stability score of the output results of different paths, performing path pruning and filtering, and outputting the intent structure vector.
[0007] In a possible implementation, a semantic evolution path graph is constructed based on speech continuity constraints, including: mapping the acoustic perturbation component, prosodic structure component, and speaker voice embedding component in the acoustic state representation tensor to different subspace dimensions of the semantic latent space, forming a hierarchically coupled acoustic-semantic joint representation structure; based on the acoustic-semantic joint representation structure, performing multi-directional semantic offset expansion around the main semantic center vector within each time window to generate a semantic candidate cluster containing the main semantic vector and semantic offset vectors, wherein the semantic offset direction is jointly determined by the acoustic perturbation trend and the prosodic change direction; performing cross-window alignment processing on the semantic candidate clusters in adjacent time windows, establishing temporal association relationships between semantic nodes based on the spatial distance between semantic vectors, the consistency of prosodic changes, and the degree of acoustic perturbation continuity, and assigning corresponding dependency weights to the temporal association relationships; constructing a directed semantic evolution path structure based on the temporal association relationships, so that multiple semantic candidate clusters form a path network structure that can be branched and merged in the time dimension; and forming a semantic evolution path graph based on the path network structure.
[0008] In a possible implementation, the determination of the main semantic center vector includes: within the current time frame, mapping the acoustic perturbation component and the prosodic structure component to the semantic latent space respectively to generate multiple initial semantic candidate vectors, and combining the speaker's voice embedding component to correct the expression habits of the multiple initial semantic candidate vectors to form a corrected candidate vector set; clustering the candidate vector set to identify regions with dense semantic distribution, and determining the geometric centroid of the regions with dense semantic distribution as candidate center vectors; comparing the candidate center vectors with the main semantic center vector of the previous time window across windows, and selecting the candidate center vector with the highest stability based on the semantic space displacement amplitude and the degree of consistency of prosodic changes to construct the main semantic center vector.
[0009] In a possible implementation, multi-channel feature decomposition is performed on the original speech signal, including: inputting the original speech signal into a multi-channel processor, performing hierarchical and segmented processing at different time scales to construct short-time transient signal sequences, mid-time speech rhythm sequences, and long-time semantic rhythm sequences respectively; performing spectral decomposition and energy perturbation separation processing on the short-time transient signal sequences to extract acoustic perturbation components reflecting frequency drift trends and local energy jump characteristics; performing beat structure analysis and stress distribution analysis on the mid-time speech rhythm sequences to extract prosodic structure components representing prosodic rhythm and the direction of intonation changes; performing speaker stability feature separation on the long-time semantic rhythm sequences to extract speaker voice embedding components reflecting speaker expression habits and vocal characteristics; performing cross-scale alignment and correlation decoupling processing on the acoustic perturbation components, prosodic structure components, and speaker voice embedding components, and performing tensor quantization reconstruction to construct an acoustic state representation tensor.
[0010] In possible implementations, different time windows are constructed using a sliding window approach, and overlapping intervals are set between adjacent time windows to maintain the continuity of the semantic candidate distribution space in the time dimension.
[0011] In a possible implementation, mapping the intent structure vector to a structured semantic output further includes: performing output verification based on the structured semantic output and establishing verification feedback; and using the verification feedback to perform iterative optimization and adjustment of speech data processing.
[0012] This application also provides a speech data processing system based on a large language model. The system includes: a speech signal feature decomposition module, used to perform multi-channel feature decomposition on the received raw speech signal to construct an acoustic state representation tensor containing acoustic perturbation components, prosodic structure components, and speaker voice embedding components; a semantic evolution path graph construction module, used to construct a semantic candidate distribution space based on the acoustic state representation tensor, generate multiple sets of semantic hypothesis vectors for speech segments in different time windows, and construct a semantic evolution path graph according to speech continuity constraints, wherein the semantic evolution path graph includes semantic nodes and temporal dependency weights between nodes; an uncertainty calculation module, used to perform uncertainty calculation on the semantic evolution path graph to generate a semantic uncertainty function representing semantic ambiguity and speech perturbation sensitivity; an intent structure vector construction module, used to dynamically construct inference depth control parameters according to the semantic uncertainty function, and input the inference depth control parameters to the multi-layer inference path scheduling unit of the large language model to perform output competition selection under the multi-path hypothesis parallel inference mechanism to construct an intent structure vector; and a semantic output mapping module, used to map the intent structure vector into structured semantic output.
[0013] The proposed speech data processing method and system based on a large language model performs multi-channel feature decomposition on the received raw speech signal to construct an acoustic state representation tensor; constructs a semantic candidate distribution space, generates multiple sets of semantic hypothesis vectors, and constructs a semantic evolution path graph; generates a semantic uncertainty function representing semantic ambiguity and speech perturbation sensitivity; dynamically constructs inference depth control parameters and inputs them into the multi-layer inference path scheduling unit of the large language model to perform output competition selection, constructs an intent structure vector, and maps it to structured semantic output. This solves the technical problems in existing technologies, such as the difficulty in accurately and effectively separating acoustic features in complex acoustic environments, resulting in low accuracy in understanding highly ambiguous speech, and the lack of dynamic inference capabilities to cope with semantic uncertainty risks. It achieves the technical effects of improving the semantic understanding accuracy and ambiguity resolution capabilities of speech interaction, and reducing the business misjudgment rate and risks. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 This is a schematic diagram of the speech data processing method based on a large language model provided in an embodiment of this application.
[0016] Figure 2 This is a schematic diagram of the structure of a speech data processing system based on a large language model, provided in an embodiment of this application.
[0017] Figure labeling: Speech signal feature decomposition module 10, semantic evolution path graph construction module 20, uncertainty calculation module 30, intent structure vector construction module 40, semantic output mapping module 50. Detailed Implementation
[0018] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on the specific implementation methods, structures, features and effects of the present invention.
[0019] This application provides a speech data processing method based on a large language model, such as... Figure 1 As shown, the method includes: Step S100: After receiving the original speech signal, perform multi-channel feature decomposition on the original speech signal to construct an acoustic state representation tensor that includes acoustic perturbation components, prosodic structure components and speaker voice embedding components.
[0020] Step S100 further includes inputting the original speech signal into a multi-channel processor, performing hierarchical and segmented processing at different time scales to construct short-term transient signal sequences, mid-term speech rhythm sequences, and long-term semantic rhythm sequences, respectively; performing spectral decomposition and energy perturbation separation processing on the short-term transient signal sequences to extract acoustic perturbation components reflecting frequency drift trends and local energy jump characteristics; performing beat structure analysis and stress distribution analysis on the mid-term speech rhythm sequences to extract prosodic structure components representing rhythmic patterns and the direction of intonation changes; performing speaker stability feature separation on the long-term semantic rhythm sequences to extract speaker voice embedding components reflecting speaker expression habits and vocal characteristics; performing cross-scale alignment and correlation decoupling processing on the acoustic perturbation components, prosodic structure components, and speaker voice embedding components, and performing tensor quantization reconstruction to construct an acoustic state representation tensor.
[0021] Preferably, the received raw speech signal is input into a multi-channel processor to perform multi-channel feature decomposition. The multi-channel processor is a signal processing unit capable of performing multiple feature extractions with different configurations. Each channel uses a different time window scale to perform parallel analysis of the speech signal, including short-time analysis channels, medium-time analysis channels, and long-time analysis channels, which are used to capture transient changes, syllable fluctuations, and global stable features of the speech signal, respectively. Specifically, hierarchical segmentation processing refers to the multi-channel processor slicing the raw speech signal according to different time window lengths to obtain short-time transient signal sequences, medium-time speech rhythm sequences, and long-time semantic rhythm sequences. The short-time transient signal sequences are sampled using extremely short time windows (e.g., 10-50 milliseconds) to capture rapidly changing physical acoustic details in the speech; the medium-time speech rhythm sequences are sampled using medium-length time windows (e.g., 100-500 milliseconds) to capture fluctuations in rhythmic units such as syllables and stresses; and the long-time semantic rhythm sequences are sampled using even longer time windows (e.g., more than 1 second or even an entire sentence) to capture macroscopic expressive rhythms such as changes in the speaker's speech rate and pause patterns.
[0022] Preferably, mathematical transformations are performed on short-time transient signal sequences, including spectral decomposition and energy perturbation separation. Specifically, the time-domain signal is converted to a frequency-domain distribution using Fourier transform, the intensity of each frequency component is calculated, and non-stationary, abrupt perturbation components are extracted from the spectrum to obtain acoustic perturbation components, representing frequency drift trends, such as continuous increases or decreases in a frequency component over time, and local energy jumps, such as sudden environmental noise spikes. Phonetic feature analysis is performed on mid-time speech rhythm sequences, including beat structure analysis and stress distribution analysis. Beat structure analysis is used to detect the periodicity of syllable or stress occurrences in speech and calculate the tempo. Stress distribution analysis is used to identify reinforced syllables, such as those with higher volume or pitch, and analyze their distribution patterns to obtain prosodic structure components, representing prosodic rhythm, such as syllable duration patterns and the direction of intonation changes, such as a rising pitch at the end of a sentence indicating a question or a falling pitch indicating a statement. Voiceprint feature analysis is performed on long-term semantic rhythm sequences, including speaker stable feature separation. This involves filtering out specific text content from long-term speech and extracting physical features that do not change with the content and belong to the speaker, such as vocal tract length and statistical distribution of vocal cord vibration fundamental frequency. This results in fixed-length speaker voice embedding components that reflect the speaker's expression habits and vocal characteristics.
[0023] Preferably, since the acoustic perturbation component, prosodic structure component, and speaker voice embedding component come from different time scales, cross-scale alignment is performed to unify them under the same time reference frame. For example, the mid-time prosodic unit and long-time speaker feature corresponding to each short-time perturbation are determined, and then correlation decoupling processing is performed to remove redundant information between the three components. For example, the speaker features may contain some prosodic information. The overlapping information is stripped by orthogonalization operation to ensure that each component represents a truly independent and complementary speech attribute. Then, the three independent components after alignment and decoupling are combined into a multi-dimensional data array according to a preset dimensional order, such as the time axis or feature type axis, and finally a structured acoustic state representation tensor is generated. The original speech signal is represented as the speaker's identity features, tone features, and environmental perturbations at a certain point in time.
[0024] Step S200: Construct a semantic candidate distribution space based on the acoustic state representation tensor, generate multiple sets of semantic hypothesis vectors for speech segments in different time windows, and construct a semantic evolution path graph based on speech continuity constraints, wherein the semantic evolution path graph includes semantic nodes and temporal dependency weights between nodes.
[0025] Preferably, the acoustic state representation tensor containing perturbation, prosody, and speaker information is used as input and mapped using a deep neural network model. Due to the inherent ambiguity of acoustic signals, such as homophones and unclear pronunciations, the model does not directly provide a unique and definite word. Instead, it calculates multiple possible semantic interpretations and their probability distributions in a high-dimensional semantic vector space, forming a semantic candidate distribution space containing various semantic possibilities. Then, speech segments within different time windows are discretely sampled from the semantic candidate distribution space. For each time window, based on the acoustic features within that time window, multiple most probable semantic points are searched in the semantic candidate distribution space. Each point is a semantic hypothesis vector, i.e., a candidate word / phrase within a time window. For example, window 1 might be "interest rate" or "profit". "Smoothness" is used in window 2, which may be "how much" or "how long," thus generating multiple sets of semantic hypothesis vectors. Speech continuity constraints are set to ensure that the semantic transition in time is smooth and logical. For example, the semantics of the previous window and the semantics of the next window must be close in the vector space or conform to grammatical rules. For example, the probability of a preposition being followed by a noun is much greater than that of a preposition being followed by a verb. Each semantic hypothesis vector is used as a semantic node. The probability of a semantic hypothesis from the previous window to the next window is calculated to obtain a time-dependent weight, which is used as the value on the edge connecting two nodes. The higher the time-dependent weight, the more reasonable the semantic transition. Finally, a directed semantic evolution path graph is generated to depict the process of semantic evolution through different branch paths when speech unfolds on the time axis.
[0026] Furthermore, step S200 also includes mapping the acoustic perturbation component, prosodic structure component, and speaker voice embedding component in the acoustic state representation tensor to different subspace dimensions of the semantic latent space, forming a hierarchically coupled acoustic-semantic joint representation structure; based on the acoustic-semantic joint representation structure, performing multi-directional semantic offset expansion around the main semantic center vector within each time window to generate a semantic candidate cluster containing the main semantic vector and semantic offset vector, wherein the semantic offset direction is jointly determined by the acoustic perturbation trend and the prosodic change direction; performing cross-window alignment processing on the semantic candidate clusters in adjacent time windows, establishing temporal association relationships between semantic nodes based on the spatial distance between semantic vectors, the consistency of prosodic changes, and the continuity of acoustic perturbations, and assigning corresponding dependency weights to the temporal association relationships; constructing a directed semantic evolution path structure based on the temporal association relationships, so that multiple semantic candidate clusters form a path network structure that can branch and merge in the time dimension; and forming a semantic evolution path graph based on the path network structure.
[0027] Preferably, the acoustic perturbation component, prosodic structure component, and speaker embedding component in the acoustic state representation tensor are mapped to different subspace dimensions of the semantic latent space, respectively. That is, the three components occupy different coordinate dimensions in the semantic latent space. For example, the semantic latent space has 100 dimensions, of which the first 20 dimensions are controlled by "acoustic perturbation", the middle 30 dimensions are controlled by "prosodic structure", and the last 50 dimensions are controlled by "speaker embedding". This hierarchical coupling describes the same speech segment, generating an acoustic-semantic joint representation structure. Here, acoustic perturbation represents the physical properties of speech, prosodic structure represents the expressive properties of speech, and speaker embedding represents the identity properties of speech. Clustering is performed based on the current original speech signal. The core semantic center vector is determined, which represents the meaning of the original speech without interference. Then, within each time window, a multi-directional semantic shift is performed around the core semantic center vector to expand the acoustic-semantic joint representation structure. The direction of semantic shift is determined by the acoustic perturbation trend and the prosodic change direction. The acoustic perturbation trend means that if the current speech has a frequency drift, such as being affected by noise, the semantics shifts towards the direction of "may have misheard". The prosodic change direction means that if the intonation rises, such as in an interrogative tone, the semantics shifts towards the direction of "interrogative sentence". This generates multiple semantic candidate clusters that form graph nodes, containing the core semantic vector and multiple semantic shift vectors.
[0028] Preferably, cross-window alignment processing is performed on semantic candidate clusters within adjacent time windows. That is, semantic candidate clusters in adjacent time windows are connected, and each candidate semantic in the previous window is paired with each candidate semantic in the next window for checking. The spatial distance between semantic vectors, the consistency of prosodic variation, and the continuity of acoustic perturbation are used to check whether the semantics are coherent. The spatial distance between semantic vectors is to check whether the distance between two candidate semantics in the vector space represents whether the semantics are coherent. The consistency of prosodic variation is to check whether the intonation fluctuation from the previous word to the next word conforms to the rules of natural language. The continuity of acoustic perturbation is to check whether the change of background noise is smooth. For example, a sudden interruption of noise indicates that there may be a sentence break. Based on the three judgment criteria, a dependency weight is calculated for each possible pair of candidate semantic connections. A high dependency weight means that the two candidate words read smoothly together, and a low dependency weight means that they read awkwardly. This is used as the temporal correlation between semantic nodes to form the connection edge of semantic nodes in the graph. By connecting multiple semantic nodes based on temporal relationships, a directed semantic evolution path structure is constructed, enabling multiple semantic candidate clusters to form a branching and merging path network structure in the temporal dimension. That is, the same semantic node may be connected to multiple different semantic nodes, indicating multiple directions of understanding. Different preceding semantic nodes may converge to the same subsequent semantic node, indicating that different paths lead to the same destination. Finally, a semantic evolution path graph is generated, which contains all possible sentence interpretations and their probabilities.
[0029] Furthermore, step S200 also includes mapping the acoustic perturbation component and the prosodic structure component to the semantic latent space within the current time frame to generate multiple initial semantic candidate vectors, and performing expression habit correction on the multiple initial semantic candidate vectors in conjunction with the speaker's voice embedding component to form a corrected candidate vector set; clustering the candidate vector set to identify semantically dense regions, and determining the geometric centroid of the semantically dense regions as candidate center vectors; comparing the candidate center vectors with the main semantic center vector of the previous time window across windows, and selecting the candidate center vector with the highest stability based on the semantic space displacement amplitude and the degree of consistency of prosodic changes to construct the main semantic center vector.
[0030] Preferably, the acoustic perturbation component representing environmental interference and the prosodic structure component representing tone and intonation are mapped to the semantic latent space respectively. Due to the different input sources, two different initial results are obtained, namely multiple initial semantic candidate vectors. For example, the perturbation component may map to "fixed deposit / live deposit", and the prosodic component may map to "regular / regular?" (interrogative tone). Then, the speaker's voice embedding component is introduced, that is, the speaker's identity features, such as a customer's habit of saying "check the balance" as "see how much is left". The multiple initial semantic candidate vectors are corrected for expression habits, and options that do not conform to the speaker's expression habits are eliminated to obtain a set of corrected candidate vectors. Each vector in the set represents a possible semantic interpretation and some impossible options caused by the speaker's habits have been eliminated.
[0031] Preferably, clustering is performed using DBSCAN clustering to perform density-based spatial clustering of the candidate vector set, identifying multiple vectors that are close together, indicating similar meanings, and identifying the region with the most clustered vectors as the semantically dense region, representing the most likely semantic range of the current speech segment. The geometric centroid of this semantically dense region is then determined by calculating the average of all candidate vectors in this region to determine the candidate center vector, representing the most likely meaning of the speech within the current window based on acoustic features. Then, the candidate center vector is compared across windows with the main semantic center vector of the previous time window to ensure the consistency of semantic understanding and avoid inconsistencies, based on the semantic space. The candidate center vector with the highest stability is selected based on the magnitude of displacement and the consistency of prosodic variation. The magnitude of semantic space displacement is calculated by measuring the distance between two center vectors in the semantic space. If the distance is too large, such as jumping suddenly from "deposit" to "withdrawal", it is considered unreasonable. The consistency of prosodic variation is judged by whether the change in intonation from the previous window to the current window conforms to the natural law. For example, the distance between a declarative sentence and an interrogative sentence is greater than that between another declarative sentence. Finally, from the multiple candidate centers of the current window, the point with the smoothest and most stable connection with history is selected as the main semantic center vector of the current time window. At the same time, the current acoustic features, speaker habits, and semantic and tonal coherence with the preceding text are also considered.
[0032] Step S200 further includes constructing different time windows using a sliding window approach and setting overlapping intervals between adjacent time windows to maintain the continuity of the semantic candidate distribution space in the time dimension.
[0033] Preferably, different time windows are constructed by sliding a "window" of a fixed length forward on the time axis with a fixed step. After each slide, the voice area covered by the window is a new processing unit. Assuming the window length is 200 milliseconds and the moving step is 50 milliseconds, the first window covers 0 - 200 ms, the second window covers 50 - 250 ms, and so on; since the window length is greater than the moving step, an overlapping interval is set between adjacent time windows to maintain the continuity of the semantic candidate distribution space in the time dimension and prevent semantic breaks caused by abrupt segmentation of the voice. For example, if the window length = the moving step, when a word is exactly cut in half, such as the "li" in "interest rate" at the end of window 1 and the "lv" at the beginning of window 2, neither window 1 nor window 2 can obtain the complete acoustic features, resulting in incorrect semantic understanding.
[0034] Step S300, perform uncertainty calculation on the semantic evolution path graph to generate a semantic uncertainty function representing semantic ambiguity and speech perturbation sensitivity.
[0035] Preferably, the uncertainty calculation is to perform mathematical statistics and entropy value calculation on the semantic evolution path graph, and evaluate the difficulty of understanding the semantics based on semantic ambiguity and speech perturbation sensitivity. Specifically, evaluate the semantic ambiguity for the branch situation in the semantic evolution path graph. If at a certain time node, the weights of multiple edges in the graph are very high. For example, the probabilities of "interest rate query" and "profit query" are both 45%, and the system cannot determine the semantic evolution path, then this uncertainty is quantified as semantic ambiguity. The higher the semantic ambiguity, the more possible interpretations there are, representing the polysemy of the language itself or the ambiguity of the recognition result; evaluate the speech perturbation sensitivity for the stability of the nodes in the semantic evolution path graph and the correlation between the nodes and the acoustic perturbation components, including backtracking to check each semantic node in the graph to determine whether it is generated by pure speech or speech severely contaminated by noise, and if the generation of a certain semantic node strongly depends on the perturbed acoustic features, that is, as long as the background noise changes, this semantic node disappears or changes, then the area corresponding to this semantic node is marked as high perturbation sensitivity. Among them, the speech perturbation sensitivity represents the quality of the current speech signal and the vulnerability of the understanding result to noise; finally, combine the calculation results of semantic ambiguity and speech perturbation sensitivity into a semantic uncertainty function that changes with time and semantic position, manifested as a mapping relationship. For example, between time points t1 and t2, the semantic ambiguity is high but the perturbation sensitivity is low, and near time point t3, the speech perturbation sensitivity is extremely high.
[0036] Step S400, dynamically construct an inference depth control parameter according to the semantic uncertainty function, and input the inference depth control parameter into the multi-layer inference path scheduling unit of the large language model to perform output competition selection under the multi-path hypothesis parallel inference mechanism, and construct an intention structure vector.
[0037] Step S400 further includes: decomposing the semantic uncertainty function into semantic ambiguity gradient components and acoustic perturbation coupling components to construct an uncertainty layering matrix; calculating the inference complexity score within the corresponding time window based on the uncertainty layering matrix, and mapping the inference complexity score to a multi-level inference depth level according to a preset segmented threshold interval; generating differentiated inference resource allocation vectors for different inference depth levels, the inference resource allocation vectors including the attention head activation ratio, the number of inference layer expansions, and the number of parallel hypothesis paths; using the differentiated inference resource allocation vectors as inference depth control parameters and inputting them into the multi-level inference path scheduling unit of the large language model; constructing a multi-level inference scheduling graph based on the multi-level inference path scheduling unit, mapping different hypothesis paths to independent inference subgraph structures, and dynamically controlling the activation range and iteration depth of each inference subgraph according to the differentiated inference resource allocation vectors; constructing a cross-path consistency tensor during multi-path hypothesis parallel inference, using the cross-path consistency tensor to jointly calculate the semantic consistency score, context continuity score, and perturbation stability score of the output results of different paths, performing path pruning and filtering, and outputting an intent structure vector.
[0038] Preferably, the semantic uncertainty function is decomposed into a semantic ambiguity gradient component and an acoustic perturbation coupling component, representing the degree of ambiguity and the degree of noise influence at the pure linguistic level, respectively. A two-dimensional uncertainty hierarchical matrix is constructed based on these two indicators, the semantic ambiguity gradient component and the acoustic perturbation coupling component. Each region in the matrix represents a combination type of uncertainty, such as high ambiguity and high perturbation, low ambiguity and high perturbation, etc. Based on the position of the current time window in the matrix, a comprehensive reasoning complexity score is calculated. Then, it is compared with a preset segmented threshold range, such as 0-3 points for simple, 4-6 points for medium, and 7-10 points for complex, and the reasoning complexity score is mapped to a multi-level reasoning depth level, such as level 1, level 2, and level 3. Next, different amounts of computing resources are allocated according to different inference depth levels, i.e., a differentiated inference resource allocation vector is generated, including the attention head activation ratio, the number of inference layer expansions, and the number of parallel hypothesis paths. The attention head activation ratio represents the proportion of attention heads used to process the current task, such as 30% attention heads for simple tasks and 100% attention heads for complex tasks. The number of inference layer expansions represents the number of iterative calculations performed in the model network, and the number of parallel hypothesis paths represents the number of different semantic understanding paths that are simultaneously retained for parallel computing. The differentiated inference resource allocation vector is then used as an inference depth control parameter and input into the multi-layer inference path scheduling unit in the large language model for allocating computing tasks.
[0039] Preferably, the multi-layer inference path scheduling unit constructs multiple parallel computational branches as a multi-layer inference scheduling graph based on task instructions. This means that each hypothesis path in the semantic evolution path graph is mapped into an independent computational subgraph structure within the large language model. The computational scale of each subgraph is precisely controlled according to the parameters in the inference resource allocation vector. For high-probability primary hypothesis paths, more attention and deeper iterations are allocated, while for low-probability secondary hypothesis paths, only a small amount of resources are allocated for shallow computation. During multi-path hypothesis parallel inference, the outputs of all parallel paths are integrated to form a cross-path consistency tensor. This cross-path consistency tensor is then used to comprehensively score the outputs of different paths, including calculating the semantic consistency score, contextual continuity score, and perturbation stability score for each path's output. Semantic consistency assesses whether the internal logic of the path is self-consistent; contextual continuity assesses whether the path result is connected to the historical dialogue; and perturbation stability assesses whether the path result is drastically changed due to noise interference. A joint score is then calculated, and low-scoring paths are pruned, retaining the semantic representation of the optimal path and outputting it as the intent structure vector.
[0040] Step S500: Map the intent structure vector to a structured semantic output.
[0041] Step S500 further includes performing output verification based on the structured semantic output and establishing verification feedback; and using the verification feedback to perform iterative optimization and adjustment of speech data processing.
[0042] Preferably, the intent structure vector is mapped to structured semantic output through a decoder, such as converting it into natural language text, "You are inquiring about the interest rate of a fixed deposit." Then, output verification is performed based on the structured semantic output, which may include business rule verification, user confirmation verification, and execution result verification. Among them, business rule verification is used to check whether the structured semantic output conforms to the banking business logic. For example, "Withdraw 1 million" requires verification of whether the account balance is sufficient. If not, the intent is not executable. User confirmation verification is used in dialogue-oriented scenarios to convert the structured semantic output into a confirmation question for the user, such as "Do you want to open a fixed deposit?". The user's affirmative or negative answer constitutes the feedback. Execution result verification refers to checking whether valid data is successfully returned after the operation is executed (such as a query operation).
[0043] Preferably, the verification results are integrated into a mathematical signal, i.e., verification feedback, and iterative optimization and adjustment of speech data processing are performed through the verification feedback. For example, if the feedback indicates that the user has corrected the misheard word "regular" as "current account", the weight of the acoustic perturbation component is adjusted retrospectively to reduce the probability of misjudgment in similar noisy environments; if a semantic path is rejected by the user, the dependency weight of that path in the semantic evolution path graph is reduced, so that its priority is reduced in similar scenarios in the future; if the verification feedback shows that a certain type of scenario often makes mistakes, the inference depth level corresponding to that type of scenario is increased, and more computing resources are allocated to it; then the correct original speech signal - structured semantic output - is used as new training data to incrementally update the large language model, thereby improving the semantic understanding accuracy of speech interaction, the ability to resolve ambiguities, and reducing the business misjudgment rate and risk.
[0044] In the above text, refer to Figure 1 A speech data processing method based on a large language model according to embodiments of the present invention has been described in detail. Next, reference will be made to... Figure 2 A speech data processing system based on a large language model according to an embodiment of the present invention is described.
[0045] The speech data processing system based on a large language model according to embodiments of the present invention addresses the technical problems in existing technologies, such as the difficulty in accurately and effectively separating acoustic features in complex acoustic environments, resulting in low accuracy in understanding highly ambiguous speech, and the lack of dynamic reasoning capabilities to cope with semantic uncertainty risks. It achieves the technical effects of improving the semantic understanding accuracy and ambiguity resolution capabilities of speech interaction, and reducing business misjudgment rates and risks. Figure 2 As shown, the speech data processing system based on the large language model includes: a speech signal feature decomposition module 10, a semantic evolution path graph construction module 20, an uncertainty calculation module 30, an intent structure vector construction module 40, and a semantic output mapping module 50.
[0046] The speech signal feature decomposition module 10 is used to perform multi-channel feature decomposition on the received original speech signal to construct an acoustic state representation tensor containing acoustic perturbation components, prosodic structure components, and speaker voice embedding components. The semantic evolution path graph construction module 20 is used to construct a semantic candidate distribution space based on the acoustic state representation tensor, generate multiple sets of semantic hypothesis vectors for speech segments in different time windows, and construct a semantic evolution path graph according to speech continuity constraints. The semantic evolution path graph includes semantic nodes and time dependency weights between nodes. The uncertainty calculation module 30 is used to perform uncertainty calculation on the semantic evolution path graph to generate a semantic uncertainty function that characterizes semantic ambiguity and speech perturbation sensitivity. The intent structure vector construction module 40 is used to dynamically construct inference depth control parameters based on the semantic uncertainty function, and input the inference depth control parameters to the multi-layer inference path scheduling unit of the large language model to perform output competition selection under the multi-path hypothesis parallel inference mechanism to construct the intent structure vector. The semantic output mapping module 50 is used to map the intent structure vector into structured semantic output.
[0047] The specific configuration of the intent structure vector construction module 40 will be described in detail below. The intent structure vector construction module 40 further includes: decomposing the semantic uncertainty function into semantic ambiguity gradient components and acoustic perturbation coupling components to construct an uncertainty layering matrix; calculating the inference complexity score within the corresponding time window based on the uncertainty layering matrix, and mapping the inference complexity score to a multi-level inference depth level according to a preset segmented threshold interval; generating differentiated inference resource allocation vectors for different inference depth levels, the inference resource allocation vectors including the attention head activation ratio, the number of inference layer expansions, and the number of parallel hypothesis paths; inputting the differentiated inference resource allocation vectors as inference depth control parameters to the multi-level inference path scheduling unit of the large language model; constructing a multi-level inference scheduling graph based on the multi-level inference path scheduling unit, mapping different hypothesis paths to independent inference subgraph structures, and dynamically controlling the activation range and iteration depth of each inference subgraph according to the differentiated inference resource allocation vectors; constructing a cross-path consistency tensor during multi-path hypothesis parallel inference, using the cross-path consistency tensor to jointly calculate the semantic consistency score, context continuity score, and perturbation stability score of the output results of different paths, performing path pruning and filtering, and outputting the intent structure vector.
[0048] The following describes in detail the specific configuration of the semantic evolution path graph construction module 20. The semantic evolution path graph construction module 20 further includes: mapping the acoustic perturbation component, prosodic structure component, and speaker voice embedding component in the acoustic state representation tensor to different subspace dimensions of the semantic latent space, forming a hierarchically coupled acoustic-semantic joint representation structure; based on the acoustic-semantic joint representation structure, performing multi-directional semantic offset expansion around the main semantic center vector within each time window to generate a semantic candidate cluster containing the main semantic vector and semantic offset vectors, wherein the semantic offset direction is jointly determined by the acoustic perturbation trend and the prosodic change direction; performing cross-window alignment processing on the semantic candidate clusters in adjacent time windows, establishing temporal association relationships between semantic nodes based on the spatial distance between semantic vectors, the consistency of prosodic changes, and the continuity of acoustic perturbations, and assigning corresponding dependency weights to the temporal association relationships; constructing a directed semantic evolution path structure based on the temporal association relationships, enabling multiple semantic candidate clusters to form a path network structure that can branch and merge in the time dimension; and forming a semantic evolution path graph based on the path network structure.
[0049] The following will describe in detail the specific configuration of the semantic evolution path graph construction module 20. The semantic evolution path graph construction module 20 further includes: within the current time frame, mapping the acoustic perturbation component and the prosodic structure component to the semantic latent space respectively, generating multiple initial semantic candidate vectors, and combining the speaker's voice embedding component to correct the expression habits of the multiple initial semantic candidate vectors, forming a corrected candidate vector set; clustering the candidate vector set to identify semantically dense regions, and determining the geometric centroid of the semantically dense regions as candidate center vectors; comparing the candidate center vectors with the main semantic center vector of the previous time window across windows, and selecting the candidate center vector with the highest stability based on the semantic space displacement amplitude and the degree of consistency of prosodic changes, thus constructing the main semantic center vector.
[0050] The specific configuration of the speech signal feature decomposition module 10 will be described in detail below. The speech signal feature decomposition module 10 further includes: inputting the original speech signal into a multi-channel processor, performing hierarchical and segmented processing at different time scales to construct short-term transient signal sequences, mid-term speech rhythm sequences, and long-term semantic rhythm sequences respectively; performing spectral decomposition and energy perturbation separation processing on the short-term transient signal sequences to extract acoustic perturbation components reflecting frequency drift trends and local energy jump characteristics; performing beat structure analysis and stress distribution analysis on the mid-term speech rhythm sequences to extract prosodic structure components representing rhythmic patterns and the direction of intonation changes; performing speaker stability feature separation on the long-term semantic rhythm sequences to extract speaker voice embedding components reflecting speaker expression habits and vocal characteristics; performing cross-scale alignment and correlation decoupling processing on the acoustic perturbation components, prosodic structure components, and speaker voice embedding components, and performing tensor quantization reconstruction to construct an acoustic state representation tensor.
[0051] The following section will describe in detail the specific configuration of the semantic evolution path graph construction module 20. The semantic evolution path graph construction module 20 further includes: constructing different time windows using a sliding window method, and setting overlapping intervals between adjacent time windows to maintain the continuity of the semantic candidate distribution space in the time dimension.
[0052] The specific configuration of the semantic output mapping module 50 will be described in detail below. The semantic output mapping module 50 further includes: performing output verification based on the structured semantic output and establishing verification feedback; and using the verification feedback to perform iterative optimization and adjustment of speech data processing.
[0053] The speech data processing system based on a large language model provided in this embodiment of the invention can execute the speech data processing method based on a large language model provided in this embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A speech data processing method based on a large language model, characterized in that, The method includes: After receiving the original speech signal, multi-channel feature decomposition is performed on the original speech signal to construct an acoustic state representation tensor containing acoustic perturbation components, prosodic structure components and speaker voice embedding components. Based on the acoustic state representation tensor, a semantic candidate distribution space is constructed, multiple sets of semantic hypothesis vectors are generated for speech segments in different time windows, and a semantic evolution path graph is constructed according to the speech continuity constraint. The semantic evolution path graph includes semantic nodes and time dependency weights between nodes. Uncertainty calculation is performed on the semantic evolution path graph to generate a semantic uncertainty function that characterizes semantic ambiguity and speech perturbation sensitivity; The inference depth control parameters are dynamically constructed based on the semantic uncertainty function, and the inference depth control parameters are input into the multi-layer inference path scheduling unit of the large language model to perform output competition selection under the multi-path hypothesis parallel inference mechanism and construct the intent structure vector. The intent structure vector is mapped to a structured semantic output.
2. The speech data processing method based on a large language model as described in claim 1, characterized in that, Construct the intent structure vector, including: The semantic uncertainty function is decomposed into semantic ambiguity gradient components and acoustic perturbation coupling components to construct an uncertainty hierarchical matrix. The inference complexity score is calculated based on the uncertainty layering matrix within the corresponding time window, and the inference complexity score is mapped to a multi-level inference depth level according to the preset segmentation threshold range. Differentiated inference resource allocation vectors are generated for different inference depth levels. The inference resource allocation vectors include the attention head activation ratio, the number of inference layer expansions, and the number of parallel hypothesis paths. The differentiated inference resource allocation vector is used as an inference depth control parameter and input into the multi-layer inference path scheduling unit of the large language model. A multi-layer inference scheduling graph is constructed based on the multi-layer inference path scheduling unit, which maps different hypothesis paths into independent inference subgraph structures, and dynamically controls the activation range and iteration depth of each inference subgraph according to the differentiated inference resource allocation vector. In the parallel reasoning process of multi-path hypothesis, a cross-path consistency tensor is constructed. The cross-path consistency tensor is used to jointly calculate the semantic consistency score, context continuity score and perturbation stability score of the output results of different paths, perform path pruning and filtering, and output the intent structure vector.
3. The speech data processing method based on a large language model as described in claim 1, characterized in that, Construct a semantic evolution path graph based on speech continuity constraints, including: The acoustic perturbation component, prosodic structure component, and speaker voice embedding component in the acoustic state representation tensor are mapped to different subspace dimensions of the semantic latent space, forming a hierarchically coupled acoustic-semantic joint representation structure. Based on the acoustic-semantic joint representation structure, a multi-directional semantic offset is expanded around the main semantic center vector within each time window to generate a semantic candidate cluster containing the main semantic vector and the semantic offset vector. The semantic offset direction is determined by the acoustic perturbation trend and the prosodic change direction. Cross-window alignment is performed on semantic candidate clusters within adjacent time windows. Temporal associations between semantic nodes are established based on spatial distance between semantic vectors, consistency of prosodic variation, and continuity of acoustic perturbation. Corresponding dependency weights are assigned to the temporal associations. Based on the aforementioned temporal correlation, a directed semantic evolution path structure is constructed, enabling multiple semantic candidate clusters to form a path network structure that can branch and merge in the temporal dimension. A semantic evolution path graph is formed based on the path network structure.
4. The speech data processing method based on a large language model as described in claim 3, characterized in that, The determination of the main semantic center vector includes: Within the current time frame, the acoustic perturbation component and the prosodic structure component are mapped to the semantic latent space respectively, generating multiple initial semantic candidate vectors. The speaker's voice embedding component is then combined to correct the expression habits of the multiple initial semantic candidate vectors, forming a corrected candidate vector set. The candidate vector set is clustered to identify semantically dense regions, and the geometric centroid of the semantically dense regions is determined as the candidate center vector. The candidate center vector is compared with the main semantic center vector of the previous time window across windows. The candidate center vector with the highest stability is selected based on the semantic space displacement amplitude and the degree of consistency of prosodic change, and the main semantic center vector is constructed.
5. The speech data processing method based on a large language model as described in claim 1, characterized in that, Perform multi-channel feature decomposition on the original speech signal, including: The original speech signal is input into a multi-channel processor and processed in a hierarchical and segmented manner at different time scales to construct short-time transient signal sequences, medium-time speech rhythm sequences, and long-time semantic rhythm sequences, respectively. The short-time transient signal sequence is subjected to spectral decomposition and energy perturbation separation processing to extract acoustic perturbation components that reflect frequency drift trends and local energy jump characteristics; We perform beat structure analysis and stress distribution analysis on the mid-time speech rhythm sequence to extract prosodic structure components that represent the rhythm and intonation direction. Speaker stable feature separation is performed on the long-term semantic rhythm sequence to extract speaker voice embedding components that reflect the speaker's expression habits and vocal characteristics; The acoustic perturbation component, prosodic structure component, and speaker voice embedding component are subjected to cross-scale alignment and correlation decoupling, and tensor quantization is performed to reconstruct the acoustic state representation tensor.
6. The speech data processing method based on a large language model as described in claim 1, characterized in that, Different time windows are constructed using a sliding window approach, and overlapping intervals are set between adjacent time windows to maintain the continuity of the semantic candidate distribution space in the time dimension.
7. The speech data processing method based on a large language model as described in claim 1, characterized in that, Mapping the intent structure vector to structured semantic output also includes: Output verification is performed based on the structured semantic output, and verification feedback is established; The verification feedback is used to perform iterative optimization and adjustment of the voice data processing.
8. A speech data processing system based on a large language model, characterized in that, The system is used to implement the speech data processing method based on a large language model as described in any one of claims 1 to 7, and the system comprises: The speech signal feature decomposition module is used to perform multi-channel feature decomposition on the original speech signal after receiving it, and to construct an acoustic state representation tensor that includes acoustic perturbation components, prosodic structure components and speaker voice embedding components. The semantic evolution path graph construction module is used to construct a semantic candidate distribution space based on the acoustic state representation tensor, generate multiple sets of semantic hypothesis vectors for speech segments in different time windows, and construct a semantic evolution path graph according to speech continuity constraints. The semantic evolution path graph includes semantic nodes and temporal dependency weights between nodes. The uncertainty calculation module is used to perform uncertainty calculation on the semantic evolution path graph and generate a semantic uncertainty function that characterizes the degree of semantic ambiguity and the sensitivity to speech perturbation. The intent structure vector construction module is used to dynamically construct inference depth control parameters based on the semantic uncertainty function, and input the inference depth control parameters into the multi-layer inference path scheduling unit of the large language model to perform output competition selection under the multi-path hypothesis parallel inference mechanism to construct the intent structure vector. The semantic output mapping module is used to map the intent structure vector into structured semantic output.