Thought chain and thought mode auxiliary speech generation method and device, equipment and medium
By using thought chain and thought modality-assisted speech generation methods, the problem of insufficient naturalness and flexibility in speech emotion expression in existing technologies is solved. It enables natural language to specify complex, diverse and subtle emotion expressions, improves the naturalness of speech synthesis and the freedom of emotion control, and is applicable to fintech and healthcare business scenarios.
Patent Information
- Application Number
- CN202511492013.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing technologies struggle to freely specify complex, diverse, and nuanced emotional expressions in speech through natural language, resulting in a lack of naturalness and flexibility in speech synthesis, which affects the service quality of financial intelligent voice systems and medical and health voice systems.
The method employs a thought chain and thought modality-assisted speech generation approach. By receiving source text and text prompts, it generates emotion control vectors, processes the source text based on the thought chain mechanism to generate phoneme sequences, processes the emotion control vectors based on the thought modality mechanism to generate audio feature sequences, performs time alignment operations, and finally generates speech waveforms.
It enables flexible specification of emotional expression in speech using natural language, improving the naturalness of speech synthesis, the subtlety of expression, and the freedom of emotional control, making it suitable for fintech and healthcare business scenarios.
Smart Images

Figure CN120977289B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, in particular to a thinking chain and thinking mode auxiliary speech generation method, device, equipment and medium. BACKGROUND
[0002] The existing emotional speech synthesis technology mainly relies on limited pre-defined emotional labels (such as "happy", "sad", "angry", etc.) or fixed acoustic control parameters (such as pitch, speech rate, energy, etc.) to adjust the emotional expression of the speech. This method has certain effect in the simple emotional switching scene, but has obvious limitations in dealing with more complex, delicate and dynamic emotional expression, resulting in lack of naturalness and expression level in the speech synthesis result.
[0003] In the field of financial technology business, with the wide deployment of intelligent customer service, voice risk control, virtual assistant and other applications, the system needs to generate speech content with real and natural emotions according to different business scenarios to improve user trust and interaction efficiency. However, due to the dependence on fixed labels or parameters control, the existing technology is difficult to generate diverse and natural emotional speech for complex financial emotional scenarios (such as calming anxious customers, explaining risk results, guiding rational decision-making), which affects the service quality of financial intelligent voice systems.
[0004] In the field of medical and health business, voice interaction is widely used in remote medical consultation, rehabilitation escort, psychological intervention and other scenarios. The system not only needs to provide semantically accurate expression, but also needs to have flexible and delicate emotional expression ability, so as to provide language feedback with temperature in the process of patient emotional fluctuation, psychological counseling and rehabilitation encouragement. However, due to the lack of free and flexible emotional control means in the existing emotional speech synthesis method, the intelligent voice system in the medical and health field is difficult to dynamically adjust the speech expression style according to the actual needs of patients, which affects the auxiliary communication and user experience of the system.
[0005] In addition, the existing method is generally difficult to realize free emotional control based on natural language, and users cannot flexibly specify the required emotional expression mode through more intuitive language prompts, which limits the applicability and interaction efficiency of the system in multi-industry and multi-task scenarios. In summary, how to break through the limitations of pre-defined labels and fixed parameters, and improve the naturalness, flexibility and control accuracy of the emotional speech synthesis system, is still a prominent problem faced by current technology. SUMMARY
[0006] The main purpose of the present application is to provide a thinking chain and thinking mode auxiliary speech generation method, device, equipment and storage medium, which aims to solve the technical problem that the existing technology cannot freely specify complex, diverse and delicate speech emotional expression through natural language, resulting in lack of naturalness and flexibility in the speech synthesis result.
[0007] To achieve the above object, the present application provides a thinking chain and thinking mode auxiliary speech generation method, comprising:
[0008] receiving source text and text prompts for specifying emotional expression;
[0009] inputting the text prompts into a language model to generate an emotional control vector through the language model;
[0010] processing the source text based on a thinking chain mechanism to generate a phoneme sequence;
[0011] processing the emotional control vector based on a thinking mode mechanism to generate an audio feature sequence;
[0012] performing time alignment operation on the phoneme sequence and the audio feature sequence to generate a time alignment sequence;
[0013] inputting the time alignment sequence into a speech decoder to generate a speech waveform.
[0014] Further, to achieve the above object, the present application provides a thinking chain and thinking mode auxiliary speech generation device, comprising:
[0015] an input analysis module for receiving source text and text prompts for specifying emotional expression;
[0016] an emotional modeling module for inputting the text prompts into a language model to generate an emotional control vector through the language model;
[0017] a phoneme generation module for processing the source text based on a thinking chain mechanism to generate a phoneme sequence;
[0018] an audio feature generation module for processing the emotional control vector based on a thinking mode mechanism to generate an audio feature sequence;
[0019] a time alignment module for performing time alignment operation on the phoneme sequence and the audio feature sequence to generate a time alignment sequence;
[0020] a speech synthesis module for inputting the time alignment sequence into a speech decoder to generate a speech waveform.
[0021] Further, to achieve the above object, the present application further provides a computer device comprising a memory, a processor and a thinking chain and thinking mode auxiliary speech generation program stored on the memory and executable on the processor, wherein the thinking chain and thinking mode auxiliary speech generation program, when executed by the processor, implements the steps of the thinking chain and thinking mode auxiliary speech generation method as described above.
[0022] Further, to achieve the above object, the application further provides a computer readable storage medium, wherein the storage medium stores a thought chain and thought mode auxiliary speech generation program, and the thought chain and thought mode auxiliary speech generation program realizes the steps of the thought chain and thought mode auxiliary speech generation method when executed by a processor.
[0023] Beneficial effects: The application relates to the technical field of speech processing, can be applied to business scenarios such as financial technology and medical health, and discloses a thought chain and thought mode auxiliary speech generation method, device, equipment and medium, which comprises the following steps: receiving a source text and a text prompt for specifying emotional expression, inputting the text prompt into a language model to generate an emotional control vector, processing the source text based on a thought chain mechanism to generate a phoneme sequence, processing the emotional control vector based on a thought mode mechanism to generate an audio feature sequence, performing time alignment operation on the phoneme sequence and the audio feature sequence to generate a time alignment sequence, inputting the time alignment sequence into a speech decoder to generate a speech waveform. The thought chain mechanism and the thought mode mechanism are combined, the limitation of traditional fixed emotional tags or preset control parameters is broken, natural language is flexibly specified to express speech emotional expression, and the naturalness of speech synthesis, the delicacy of expression and the freedom of emotional control are improved. BRIEF DESCRIPTION OF DRAWINGS
[0024] The application will be further described below in combination with the drawings and embodiments, wherein:
[0025] Figure 1 An application environment schematic diagram of the thought chain and thought mode auxiliary speech generation method in an embodiment of the application;
[0026] Figure 2 A flowchart of the thought chain and thought mode auxiliary speech generation method in an embodiment of the application;
[0027] Figure 3 A functional module schematic diagram of the thought chain and thought mode auxiliary speech generation device in a preferred embodiment of the application;
[0028] Figure 4 A structure schematic diagram of a computer device in an embodiment of the application;
[0029] Figure 5 Another structure schematic diagram of a computer device in an embodiment of the application. DETAILED DESCRIPTION
[0030] It should be understood that the specific embodiments described herein are merely intended to explain the application, and are not intended to limit the application.
[0031] The thought chain and thought mode auxiliary speech generation method provided by the embodiments of the application can be applied to the business scenarios such as Figure 1In an application environment of the present application, a user terminal communicates with a server terminal through a network. The server terminal can receive source text and text prompts for specifying emotional expression from the user terminal, input the text prompts into a language model to generate an emotional control vector, process the source text based on a thinking chain mechanism to generate a phoneme sequence, process the emotional control vector based on a thinking mode mechanism to generate an audio feature sequence, perform time alignment operation on the phoneme sequence and the audio feature sequence to generate a time-aligned sequence, input the time-aligned sequence into a speech decoder to generate a speech waveform. By combining the thinking chain mechanism and the thinking mode mechanism, the present application breaks the limitations of traditional methods based on fixed emotional labels or preset control parameters, realizes flexible specification of speech emotional expression in natural language, and improves the naturalness of speech synthesis, the delicacy of expression, and the degree of freedom of emotional control. The user terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and portable wearable devices. The server terminal can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail below through specific embodiments.
[0032] Please refer to Figure 2 , Figure 2 The flowchart of an embodiment of the thinking chain and thinking mode assisted speech generation method provided by the present application is shown. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that shown here.
[0033] As Figure 2 shown, the thinking chain and thinking mode assisted speech generation method provided by the present application includes the following steps:
[0034] S10, receiving source text and text prompts for specifying emotional expression;
[0035] In this embodiment, the process of receiving source text and text prompts for specifying emotional expression includes two data acquisition links. The source text refers to the data content that needs to be synthesized into speech, which can be user input text information, system called text data, or externally interfaced structured or unstructured text content. The text prompts for specifying emotional expression are instruction information in natural language form, which have the function of expressing the emotional type, emotional tendency or expression style of the speech generated by the system as desired by the user, specifically including but not limited to emotional category words, descriptive sentences or context supplement information.
[0036] In the implementation process, first, the text information input stream is obtained through the input end, including source text data and text prompt data. The input stream can receive user input through a human-computer interaction terminal, or receive incoming text information from an external system through a network interface. The received source text is subjected to semantic analysis, and the basic content structure of the text is extracted to ensure the semantic integrity and contextual coherence of the content, facilitating subsequent processing. The text prompt information is analyzed through semantic recognition and intent analysis to analyze the emotional expression tendency, emotional attribute, or expression style requirement in the text, forming emotional expression target data to facilitate subsequent control of the emotional expression process when generating speech.
[0037] The text prompt is not limited to predefined standard emotional categories. Users can use natural language to flexibly describe emotions or expression requirements, such as inputting "please say the following content in a gentle tone," "express happy emotions," or "express with a firm voice." The system converts unstructured natural language text into structured emotional expression requirement information through natural language understanding capabilities.
[0038] Text input information can be received based on a text input module provided by an intelligent terminal device. Source text can come from user-entered paragraph text, real-time entered text content, or business data text called from existing databases or information management systems. The text prompt part can come from user-added natural language instructions, or from automatically generated expression style description information from the business system, or even from historical data extracted voice expression habit parameters to construct text prompt content.
[0039] More rich text prompt information can also be obtained through multiple rounds of interaction. Users can input text prompts for the first time, and the system returns intermediate results to guide users to supplement more specific expression styles or emotional details, improving the expression effect of the final generated speech. The input format of the source text can support multiple encoding standards, such as UTF-8, GBK, or other text encoding formats, ensuring that the system has stable text access capabilities in different language environments.
[0040] Example: In the medical health business field, the source text can be the diagnosis and treatment suggestion content of the doctor and patient communication, and the text prompt can be "please express in a soothing and gentle tone." After the system receives the above information, it generates a voice output with soothing properties through emotional understanding and expression control, which helps to relieve the patient's nervous emotions and optimizes the medical service experience.
[0041] In the field of financial technology business, the source text can be a financial product introduction, compliance prompt or customer notification information, and the text prompt can be "please express in a formal and reliable tone". The system combines this expression requirement when generating speech and outputs speech information with authority and trust, improving the professionalism of financial product promotion and the trust of customers.
[0042] The embodiment can introduce more flexible and fine emotional control information in the content generation process by receiving the source text and the text prompt for specifying emotional expression, avoiding the expression limitations caused by relying on preset labels or fixed parameters in traditional systems. Combined with the semantic information provided by the text prompt, the emotional tendency and style of subsequent content generation can be effectively guided, the naturalness and diversity of speech output can be improved, and more personalized expression that meets user needs can be achieved.
[0043] S20, input the text prompt into the language model, and generate an emotional control vector through the language model;
[0044] In this embodiment, the text prompt refers to the input information expressed by the user in natural language form for specifying the direction of speech emotional expression. This input information is not limited to standardized labels or preset categories, and can include emotional tendencies, tone styles, expression scenarios or specific semantic guidance, such as "please express in a gentle tone" or "simulate the style of financial announcements". The source of the text prompt can be user input, system-generated or transferred from other application interfaces. The language model is a deep neural network structure trained based on large-scale data, with understanding, expression and reasoning capabilities. Its internal structure can use a transformer network architecture, composed of multiple layers of self-attention mechanisms and feedforward networks. The process of inputting the text prompt into the language model is to first perform necessary preprocessing operations on the text prompt, including character encoding conversion, text standardization, word segmentation, word embedding mapping, etc. Then, the processed sequence data is input into the language model. The language model accepts the sequence at the input end and extracts context information, semantic association and potential emotional features through multiple internal networks. Finally, it outputs hidden state information or semantic representation. The emotional control vector is a low-dimensional feature representation form output by the language model. This vector compresses the emotional intent, expression style and context feature information in the text prompt, facilitating the generation and adjustment of audio data or other expression data in downstream modules. The specific form of the emotional control vector can be a fixed-dimensional real vector, a normalized probability distribution or a feature matrix transformed by encoding, and the numerical structure is determined according to the overall architecture of the system.
[0045] The processing manner of the text prompt can be adjusted based on different task requirements and application environments. For example, in the medical health field, a plurality of typical prompt templates can be preset for medical dialogues, rehabilitation guidance and the like, and the user supplements specific emotional or tone requirements through natural language. In the financial technology field, the user can input text information with a more formal, serious or specific risk prompt tendency in combination with the needs of financial broadcasting and compliance preaching. Different model versions with different scales and parameter structures can be selected for the language model to meet the needs of edge devices, cloud platforms or distributed deployment. The model can be trained based on Chinese, English or multi-lingual environment and has multi-scene adaptation capability. The generation form of the emotion control vector can adopt high-dimensional vector projection, encoding matrix compression or multi-layer aggregation mechanism to guarantee the expression ability of the output and the compatibility of the downstream module.
[0046] In this embodiment, the text prompt is input into the language model to generate an emotion control vector, which can fully excavate the semantic intention and emotional tendency in the user input, break through the limitations of traditional speech synthesis systems relying on fixed labels or preset parameters, improve the flexibility and naturalness of the generated content, make the output speech more accurately reflect the user's emotional expression needs, and enhance the personalized customization capability and the ability to adapt to diversified application scenarios of the system.
[0047] S30, processing the source text based on a chain-of-thought mechanism to generate a phoneme sequence;
[0048] In this embodiment, the source text refers to the text information input by the user or received by the system, which can be natural language expression, standardized terminology, domain-specific content or other symbolic language units. The source of the source text can include a human-computer interaction interface, an external system data interface or an automatic generation module.
[0049] The chain-of-thought mechanism is a processing manner based on multi-layer information progressive reasoning and semantic extension. By simulating the step-by-step deduction of context logic, semantic connection and expression structure in human language understanding process, more complex and coherent language information conversion is realized.
[0050] The essence of this mechanism is to gradually build a chain of thought from shallow to deep and layer by layer by introducing explicit intermediate reasoning steps in the information processing flow, so as to realize more complex semantic understanding, logical reasoning and expression generation. In this mechanism, the model generates an output chain containing intermediate reasoning steps guided by prompt examples, effectively improving the accuracy and stability in mathematical reasoning, logical question answering, complex text understanding and the like. The chain-of-thought mechanism is not limited to language model reasoning tasks, but also has theoretical and application basis in the fields of visual understanding, speech synthesis, decision system and the like.
[0051] In the field of text-to-speech (TTS), traditional methods usually adopt an end-to-end or step-by-step hierarchical conversion structure, which lacks explicit intermediate semantic link modeling, easily leading to information loss, logical disorder or unnatural expression. The introduction of the thinking chain mechanism is essentially inspired by the multi-step reasoning concept. In the mapping process from source text to phoneme sequence, it constructs a multi-layer semantic chain, dynamically optimizes the semantic integrity and expression logic of the voice content through hierarchical recursion, logical supplementation and context association. For example:
[0052] In the semantic analysis stage, the thinking chain mechanism identifies the basic facts, inference relationships and expression intentions in the text step by step.
[0053] In the phoneme mapping stage, the thinking chain mechanism dynamically adjusts the phoneme structure based on the semantic chain to ensure accurate pronunciation and natural expression.
[0054] In multi-round expression generation, the thinking chain mechanism maintains the consistency of context logic through chain information transmission, avoiding incoherent generated content.
[0055] Specifically, the thinking chain mechanism includes the following operations: first, perform structural analysis on the source text to divide the input text into an ordered sequence of language units, which can be words, phrases or syntactic structures; then, based on a multi-layer recursive analysis network, extract the context dependency and semantic logic of the source text step by-step to construct an internal thinking chain expression structure; based on the thinking chain expression, perform phoneme mapping for each language unit to convert it into corresponding phoneme symbols, which are the smallest pronunciation units in language expression, usually represented by international phonetic alphabet, specific voice coding or system-defined standard; after generating the initial phoneme symbol sequence, continue to perform progressive optimization on the phoneme sequence based on the thinking chain mechanism to identify ambiguous, ambiguous or complex pronunciation patterns in the language structure, and perform multiple iterations of extension and adjustment to finally generate a stable, complete and standardized phoneme sequence. The phoneme sequence is the basic input data for the subsequent voice generation process, determining the pronunciation content, rhythm structure and expression clarity of the generated voice.
[0056] The processing mode of the source text can be flexibly adjusted according to different language environments, industry fields or user needs. For example, in the medical and health field, the source text can include medical terms, rehabilitation guidance text or clinical communication information. The thinking chain mechanism can combine medical corpus and industry specifications to ensure accurate transcription and pronunciation of professional vocabulary. In the financial technology field, the source text can involve financial news, financial announcements or risk warning texts. The thinking chain mechanism can combine the language style and compliance requirements of the financial industry to improve the rigor and professionalism of expression. The phoneme mapping operation can be based on an international standard phonetic system or a custom phoneme library according to the pronunciation rules of a specific language. The multi-layer reasoning network of the thinking chain mechanism can use recurrent neural networks, transformer structures or other deep learning models to enhance the accuracy and logical coherence of expression under complex language structures.
[0057] By processing the source text based on the thinking chain mechanism and generating phoneme sequences, the embodiment can effectively improve the quality of text-to-phoneme conversion, avoid ambiguity, omission or unnatural expression problems under traditional rule mapping methods, enhance the coherence of language expression structure and the accuracy of speech synthesis, and achieve more natural, clear and user-expected speech output, especially suitable for application scenarios in medical and health and financial technology fields that have high requirements for professional terms, semantic logic and expression specification.
[0058] S40, processing the emotion control vector based on a thinking mode mechanism to generate an audio feature sequence;
[0059] In this embodiment, the thinking mode mechanism is a structured processing method that simulates human multi-modal information fusion and thinking conversion rules to improve the consistency and flexibility of the system in complex expression, semantic control and cross-modal feature generation. In this processing process, the emotion control vector, as a high-dimensional feature set representing the user's emotional expression intention, usually comes from natural language prompts, context information or external environment input. The thinking mode mechanism effectively maps the emotional information in the vector into an audio feature sequence with time continuity and expression logic through multi-layer information conversion, feature fusion and time series modeling.
[0060] In actual implementation, the thinking mode mechanism generally includes the following core steps:
[0061] First, perform cross-modal alignment operation on the emotion control vector. This process converts the emotional semantic information in the language space into the basic representation in the acoustic feature space based on a specific modal mapping network, generating an acoustic feature base vector. The base vector carries the initial mapping result of the user's emotional intention in the acoustic domain, has expression directionality but lacks specific time dynamic structure.
[0062] Secondly, based on the thinking mode mechanism, the acoustic feature basis vector is fused with the time state code to generate an initial acoustic feature frame. The time state code is derived from the preset or dynamically updated time structure information, such as expression rhythm, speech rate change, prosodic contour, etc. In the fusion process, through multi-layer nonlinear mapping and sequence modeling, it is ensured that the initial acoustic feature frame output has basic time logic and rhythm characteristics.
[0063] Further, based on the thinking mode mechanism, the initial acoustic feature frame is taken as the starting point to iteratively perform the time sequence evolution operation to generate an acoustic feature frame sequence. This operation simulates the dynamic evolution process of thinking and emotion in human expression, and through recursive structure, autoregressive mechanism or variational inference model, the time structure of the acoustic feature frame is gradually expanded, the coherence and dynamic change of multi-step expression logic are realized, and an acoustic feature frame sequence with complete time structure is generated.
[0064] Finally, the acoustic feature frame sequence is subjected to feature normalization operation to generate an audio feature sequence. The operation includes amplitude normalization, dynamic range compression, spectral smoothing or other acoustic feature optimization processing, to ensure that the generated audio feature sequence meets the input requirements of the subsequent speech decoder, and at the same time, to improve the consistency and naturalness of the expression effect.
[0065] Overall, the thinking mode mechanism converts the static emotion control vector into a dynamic, coherent and natural audio feature sequence through a series of structured operations described above, breaking through the technical bottlenecks of limited emotional expression, missing time structure and insufficient expression delicacy in traditional systems, and has wide applicability and expansion space in speech synthesis, speech expression control and other scenarios.
[0066] By inputting the emotion control vector into the thinking mode mechanism and combining cross-modal alignment, time state code fusion, time sequence evolution and feature normalization, etc., the embodiment can effectively improve the expression accuracy and dynamic coherence of emotional information in the audio feature generation process, avoid the problems of lack of delicate expression of emotional information, unnatural time structure and incoherent emotion transition in traditional systems, and further enhance the naturalness, flexibility and user perception consistency of the speech generation result in complex expression scenarios.
[0067] S50, performing time alignment operation on the phoneme sequence and the audio feature sequence to generate a time alignment sequence;
[0068] In this embodiment, the phoneme sequence is a symbol sequence representing speech pronunciation information generated based on the source text, usually composed of a series of phoneme symbols, which can be represented by the International Phonetic Alphabet or other standard phoneme sets defined in language models. The audio feature sequence is a multi-dimensional data sequence related to speech features generated based on the emotion control vector. The audio features usually include mel-frequency spectrum, mel-frequency cepstral coefficients, pitch, energy, speech rate information, and other data structures reflecting sound properties.
[0069] The time alignment operation refers to the synchronization mapping of each phoneme symbol in the phoneme sequence and the feature frame in the audio feature sequence in the time dimension. The phoneme symbol itself does not have specific time information, so it is necessary to combine the duration information of the audio feature sequence to realize the time synchronization of the phoneme sequence and the audio feature sequence by calculating their corresponding relationship in the time dimension, ensuring that the final generated speech content is consistent with the language content and speech features in the time structure.
[0070] The time alignment sequence is a structured data sequence generated based on the time mapping relationship between the phoneme sequence and the audio feature sequence, usually including phoneme symbols, corresponding audio feature segments, time start information, time end information, etc. This sequence is used to guide the subsequent speech decoding process to ensure that the generated speech waveform strictly corresponds to the phoneme information in time, improving the coherence and expression accuracy of the speech.
[0071] The time alignment of the phoneme sequence and the audio feature sequence can be achieved in the following way. First, calculate the theoretical duration information of each phoneme symbol in the phoneme sequence. This duration information can be obtained through statistical data, language rules, neural network prediction models, etc. Then calculate the overall duration of the audio feature sequence, which is usually determined by the number of feature frames and the single-frame duration parameter.
[0072] If the overall duration of the audio feature sequence is less than the theoretical total duration of the phoneme sequence, the audio feature sequence can be expanded by interpolation methods, including linear interpolation, polynomial interpolation, neural network-based feature expansion, etc., to generate an adjusted audio feature sequence. Then, based on the duration information of the phoneme sequence, the adjusted audio feature sequence is divided into several audio feature segments, the number of segments is the same as the number of phoneme symbols, and the division position corresponds to the duration information.
[0073] If the theoretical total duration of the phoneme sequence is less than the overall duration of the audio feature sequence, the duration information of the phoneme sequence can be adjusted by the equal proportion duration expansion method to make it consistent with the overall duration of the audio feature sequence. Then, divide the audio feature sequence according to the adjusted phoneme duration information to generate the audio feature segment sequence.
[0074] Finally, each phoneme symbol is bound to the corresponding audio feature segment, time start and end information is added, and it is combined into a time-aligned sequence in chronological order for subsequent speech decoding.
[0075] Example: In the medical health business field, for the doctor-patient communication voice system, doctors can input medical text information, and the system synchronizes medical terminology with voice features based on time alignment operations, so that the synthesized voice accurately restores medical professional terms in time expression, avoiding information miscommunication caused by chaotic pronunciation rhythm.
[0076] In the financial technology business field, for intelligent customer service and financial broadcast systems, time alignment operations ensure that financial terminology, amount information, risk prompts, and other voice content are strictly corresponding, improving the rigor and accuracy of voice expression, reducing information ambiguity in financial communication, and ensuring the reliability of business process compliance and user experience.
[0077] The above time alignment operation establishes a precise correspondence between the phoneme sequence and the audio feature sequence in the time dimension, effectively solving the problems of inaccurate pronunciation, chaotic pronunciation rhythm, and semantic expression errors caused by misalignment of phoneme information and audio features in traditional speech synthesis. The unity of voice content in time structure and language content, audio features is achieved, improving the naturalness, accuracy, and coherence of voice output.
[0078] S60, input the time-aligned sequence into the speech decoder to generate a speech waveform.
[0079] In this embodiment, the time-aligned sequence is a data collection generated by combining the phoneme sequence and the audio feature sequence through time mapping and synchronization operations, containing phoneme information, corresponding audio feature segments, and time marker data. The speech decoder is a module or system that can restore structured voice feature information to continuous speech waveform, usually based on neural network model implementation, common architectures include convolutional neural network-based decoding structure, autoregressive structure-based decoder, or diffusion modeling, and waveform generation system based on generative adversarial network.
[0080] Inputting the time-aligned sequence into the speech decoder means extracting audio feature information from the time-aligned sequence, arranging it into the required data format for the decoder in chronological order, and inputting it into the processing unit of the speech decoder. The speech decoder reconstructs the speech signal frame by frame or step by step based on the input time sequence information, and outputs audible continuous speech waveform through feature restoration, signal decoding, parameter mapping, waveform generation, etc.
[0081] The generating speech waveform refers to the process of converting the feature information in the time-aligned sequence into audio data conforming to the auditory standard through the processing of the speech decoder, and the audio data is usually in a standard waveform format, such as PCM (Pulse Code Modulation) data or other common speech waveform data formats, and the output result can be directly played or used for subsequent audio processing.
[0082] The input of the time-aligned sequence into the speech decoder and the generation of the speech waveform can be realized based on the following manner. First, all audio feature segments are extracted from the time-aligned sequence, reorganized into a continuous audio feature sequence according to the time label information, arranged into a matrix or tensor structure, and input into the speech decoder.
[0083] The speech decoder can adopt an architecture based on an autoregressive structure, input audio feature information frame by frame, output waveform sample data corresponding to the current frame each time, and the output of the current frame can be used as part of the condition for the next frame input, and the whole waveform data generation is completed through iterative loop.
[0084] In the specific design of the speech decoder, it usually includes a multi-layer neural network structure, a feature extraction unit, a feature mapping module, a time sequence convolution structure, a waveform reconstruction module, etc., which converts the input audio feature sequence into audio signal output through nonlinear mapping and feature fusion technology.
[0085] Through the above steps, the embodiment can accurately restore the audio feature information in the time-aligned sequence to the speech waveform, effectively solve the problems of speech blur, breakage and distortion caused by feature distortion, timing error or insufficient restoration precision in traditional speech generation, improve the naturalness, clarity and intelligibility of speech synthesis, and ensure the real restoration of language expression content in the speech output result.
[0086] The present application relates to the technical field of speech processing, which can be applied to financial technology and medical health business scenarios, and discloses a thought chain and thought mode auxiliary speech generation method, device, equipment and medium, which comprises the following steps: receiving a source text and a text prompt for specifying emotional expression, inputting the text prompt into a language model to generate an emotional control vector, processing the source text based on a thought chain mechanism to generate a phoneme sequence, processing the emotional control vector based on a thought mode mechanism to generate an audio feature sequence, performing a time alignment operation on the phoneme sequence and the audio feature sequence to generate a time-aligned sequence, inputting the time-aligned sequence into a speech decoder to generate a speech waveform. The present application breaks the limitation of traditional fixed emotional label or preset control parameter by combining the thought chain mechanism and the thought mode mechanism, realizes flexible specification of speech emotional expression in natural language, and improves the naturalness of speech synthesis, the delicacy of expression, and the freedom of emotional control.
[0087] In one embodiment, the step S20 described above comprises:
[0088] S201, performing word segmentation processing on the text prompt to generate a word sequence;
[0089] S202, converting the word sequence into a word embedding vector sequence;
[0090] S203, inputting the word embedding vector sequence into a language model to generate a hidden state sequence through an encoding layer of the language model;
[0091] S204, performing a global pooling operation on the hidden state sequence to generate a sentiment control vector.
[0092] In the present embodiment, the text prompt refers to natural language text information input by a user to express an intended speech emotion or tone style, and the source can include a short sentence, a keyword group, a complete sentence, or a paragraph. The text prompt is usually used to guide the system to generate speech content with a specific tone, intonation, and emotional tendency, and has content openness and expression diversity.
[0093] Word segmentation processing refers to an operation of dividing continuous text information into word-level units according to linguistic rules for the language structure of the text prompt, and is suitable for segmentation methods in different language environments. In a Chinese scenario, word segmentation can be performed based on dictionary matching, statistical models, or deep learning methods, and in an English or pinyin language environment, boundaries can be delimited by spaces and punctuation marks. The generated word sequence is a structured language unit set arranged in text order, facilitating subsequent vectorization operations.
[0094] Converting the word sequence into a word embedding vector sequence refers to mapping each word to a fixed-dimensional dense vector representation through table lookup mapping, neural network encoding, or model parameter generation. The word embedding vector sequence retains the language structure and semantic feature information of the text prompt and has a numerical form, facilitating input into a neural network structure.
[0095] The encoding layer of the language model is usually a sequence feature extraction unit based on a deep neural network, and the specific form can include a multi-layer Transformer encoding structure, a recurrent neural network layer, a convolutional neural network stack, or a hybrid structure. The encoding layer captures the context association and sequence information modeling through multi-layer feature transformation of the word embedding vector sequence, generates a hidden state sequence, and the hidden state sequence embodies the high-dimensional expression capability of the text prompt in the semantic space, containing complete text context information.
[0096] The global pooling operation refers to a whole information compression process for the hidden state sequence, and specific manners include average pooling, maximum pooling, attention weighted pooling, sequence head and tail extraction, etc. Through the process, the multi-dimensional and multi-step hidden state sequence is integrated into a fixed length vector expression. The generated emotion control vector is a high-order abstract representation of the semantic features and expression intention of the text prompt, has stability, compactness and expression accuracy, and is convenient for use as a control parameter in subsequent speech generation.
[0097] The embodiment can realize effective conversion from natural language text prompt to high-dimensional emotion control vector through the above operations, and solve the problems of single emotion expression and control parameter dependence on fixed labels in existing systems. The process retains the semantic richness and expression freedom of the text prompt, and the generated emotion control vector has efficient expression ability for diversified and delicate emotions, improves the flexibility and naturalness of emotion control in the speech synthesis process, and enhances the response effect of the system to complex emotional instructions.
[0098] In one embodiment, the above step S30 comprises:
[0099] S301, the source text is segmented into an ordered word unit sequence;
[0100] S302, performing phoneme symbol mapping operation on the ordered word unit sequence to generate an initial phoneme symbol sequence;
[0101] S303, based on the ordered word unit corresponding to each phoneme in the initial phoneme symbol sequence, analyzing the semantic features of the phoneme in the context;
[0102] S304, based on the semantic features and the phoneme coarticulation strategy between adjacent phonemes, positioning the position where the connecting phoneme needs to be inserted, and inserting the connecting phoneme at the corresponding position to generate a sequence after inserting the connecting phoneme;
[0103] S305, based on the emotion state classification result or the intonation turning point of the context semantics of the phoneme corresponding ordered word unit, inserting prosodic emphasis markers at the position where the stress probability is greater than the preset threshold or the syntactic boundary, to generate a sequence after inserting prosodic markers;
[0104] S306, taking the sequence after inserting prosodic markers as a new current phoneme sequence, iteratively performing semantic feature analysis operation, connecting phoneme insertion operation and prosodic emphasis marker insertion operation, taking the phoneme sequence generated in the last iteration as input, updating the context information and continuing processing until covering all ordered word units corresponding to the phonemes in the initial phoneme symbol sequence, to generate an extended phoneme sequence;
[0105] S307, performing phoneme boundary optimization operation on the extended phoneme sequence to generate a phoneme sequence.
[0106] In this embodiment, the source text refers to the complete semantic information input by the user for voice generation, usually in the form of natural language, including vocabulary, phrases, sentences or multi-sentence combinations, and the source is not limited to user input, text interface call or other system transmission, with semantic integrity and clear expression. The source text is divided into an ordered word unit sequence, which means that the source text is divided into several smallest language units with semantic independence according to language rules, punctuation, space division or specific word segmentation algorithm, and the arrangement order in the original text is kept. The generated word unit sequence retains the language structure of the text, facilitating subsequent voice information processing.
[0107] Performing phoneme symbol mapping operation on the ordered word unit sequence means converting each word unit into corresponding phoneme symbol according to language pronunciation rules or phonetic mapping table. Phoneme symbol is a standardized symbol set representing the smallest unit of speech in language, usually referring to International Phonetic Alphabet (IPA) or phoneme system of specific language, and the mapping operation can be based on static mapping dictionary, statistical pronunciation model or deep learning model. The initial phoneme symbol sequence is the phonetic expression result of the original word unit after conversion, which embodies the basic pronunciation information of the source text.
[0108] Performing step-by-step expansion operation on the initial phoneme symbol sequence based on the thought chain mechanism means dynamically supplementing or adjusting the structure content in the phoneme symbol sequence through introducing multi-level context understanding and semantic association process. The thought chain mechanism is derived from the simulation of human thinking process, usually implemented through multi-layer recursive structure of neural network, sequence-to-sequence modeling, context attention mechanism or logical reasoning module. The expansion operation includes but is not limited to inserting necessary prosodic phonemes, filling in speech connectors, adjusting phoneme expression according to context, etc. The expansion of phoneme sequence further optimizes the fluency and naturalness of pronunciation on the basis of preserving the original semantic information.
[0109] In building a system with language understanding depth and voice generation performance, phoneme sequence is no longer a simple combination of pronunciation instructions, but a plastic expression unit carrying dynamic semantic structure and auditory performance intention. The initial phoneme symbol sequence not only maps the linear expression of the text, but also implies the non-explicit emotional state, pragmatic function and auditory expectation structure stimulated by the context semantics. Therefore, expanding the semantic feature analysis of phonemes in the current context based on the ordered word unit corresponding to each phoneme is essentially a local semantic decompression of language information - it accurately transfers the meaning reconstruction task that should be undertaken by morphology and syntax to the phoneme level, forming a micro-granularity distribution of semantic tension. This distribution is not linear, but constitutes a tension field for voice synthesis, in which the structural main line of language content and the center of expression intention interact and gradually focus on the position of pronunciation action.
[0110] The insertion of connecting phones is not a modification of the "formal aesthetics" of speech, but a physical reconstruction of the rhythmic flow and the rhythmic break points in language expression. When the pronunciation characteristics of adjacent phones have discontinuity in the acoustic dimension, or show a trend of jumping from low intensity to high intensity in the semantic dimension, connecting phones are introduced as a kind of "rhythmic glide". The determination of the insertion position needs to consider both semantic features and coordinated pronunciation strategies: the former determines where the pronunciation needs to be emphasized or transitioned, and the latter restricts the form of connecting phones so that it does not destroy the language legality of phone combination and maximally relieves the impedance mismatch between phones. In real speech synthesis, the result of this insertion behavior is not only the improvement of "smooth listening", but also the provision of a continuous carrier for the subsequent prosodic contour construction, so that speech is no longer presented as a linear stack, but as a dynamic channel with internal tension and energy trend.
[0111] The introduction of emotional state and prosodic structure marks the transition of speech synthesis from "accurate expression" to "performance enhancement". The core value of language is not just to convey symbolic meaning, but how to express position, strengthen emotion, regulate rhythm and guide attention in interaction. The insertion of prosodic markers is the engineering response of speech generation system to this goal. By establishing the semantic mapping relationship between phones and corresponding word units, combined with emotion classification model and prosodic structure analysis tool, the system can identify "emotion-intensive points" and "structure turning points" in language. These points are located at positions where emotional intensity changes or semantic logic reverses, and naturally have the tendency to be emphasized in pronunciation. When the probability of stress at these positions exceeds the preset threshold, or is located at the explicit syntactic boundary, the system inserts specific prosodic markers to control the rhythm peaks and valleys of speech, so that the audience can accurately grasp the semantic core without relying on vision and context. This operation converts the "auditory labeling" of language content into the fine-tuning behavior of acoustic parameters, completing the closed loop from "cognitive intention" to "expression behavior".
[0112] On the basis of the above processing, the structural stability of the synthesized speech has not yet been achieved. The insertion of connected phonemes and prosodic markers can change the coupling relationship between the local semantic structure and the acoustic structure, and thus can break the original dependency chain or rhythm perception pattern. Therefore, the system needs to re-input the current phoneme sequence, and iteratively perform the whole process of semantic feature analysis, connected phoneme insertion, and prosodic marker injection. Each iteration is not only a readjustment of the physical structure, but also a reconstruction of the overall semantic tension field. This cycle is not a mechanical repetition, but a gradual updating of the context, forming a gradual dynamic optimization mechanism for speech generation. The goal of this mechanism is to ensure that each processed phoneme is accurately embedded in its corresponding semantic unit and obtains the optimal collaborative expression position in the pronunciation chain. Finally, the iteration will converge after all the word units covered by the phonemes reach the expression target, thereby generating a set of extended phoneme sequences with semantic tension continuity, pronunciation collaboration, and rhythm prosody.
[0113] This extended sequence itself may still have inconsistencies in physical boundaries, such as unreasonable intervals between phonemes, excessive fluctuations in pronunciation length, mismatched stress positions and rhythm benchmarks, etc. Therefore, in the final stage, the system introduces a phoneme boundary optimization mechanism to repair boundary overlaps and rhythm inconsistencies by making fine-grained adjustments to the start and end points of phonemes. The phoneme boundary optimization operation on the extended phoneme sequence refers to the accurate definition of the start and end boundary times or position relationships of each phoneme unit in the phoneme sequence, ensuring natural phoneme connection and smooth speech flow during pronunciation. The optimization operation can be combined with language model prediction, acoustic model assistance, phoneme duration statistics, or rule base. The standard phoneme sequence is the final phoneme expression form after boundary optimization, with clear, coherent, and natural language pronunciation structure.
[0114] The above steps can automatically complete phoneme-level pronunciation information extraction, speech structure expansion, and boundary adjustment while preserving the original semantic information of the source text, breaking through the problems of lack of context understanding, single pronunciation structure, and rough phoneme boundary processing in traditional phoneme mapping methods. The standard phoneme sequence has high language pronunciation characteristics, natural and coherent speech structure expression ability, improves the compatibility of language and speech in the speech generation process, and achieves more natural, accurate, and smooth speech synthesis effect, meeting the high-quality pronunciation needs in complex contexts and diverse expression scenarios.
[0115] In one embodiment, the above step S40 includes:
[0116] S401, performing a cross-modal alignment operation on the emotion control vector to generate an acoustic feature base vector;
[0117] S402, fuse the acoustic feature basis vector and the time state code based on the thinking mode mechanism to generate an initial acoustic feature frame;
[0118] S403, based on the initial acoustic feature frame, iteratively perform a time evolution operation based on the thinking mode mechanism to generate an acoustic feature frame sequence;
[0119] S404, perform feature normalization processing on the acoustic feature frame sequence to generate an audio feature sequence.
[0120] In this embodiment, the emotion control vector refers to a numerical vector expressing emotional information generated based on a text prompt, usually with multi-dimensional characteristics, used to depict the emotional tendency, tone style or semantic color of the language content expected to express, and the source includes but is not limited to the hidden state global pooling result output by the language model or other semantic mapping structure. The cross-modal alignment operation on the emotion control vector refers to establishing a mapping relationship between the emotion control vector and the target speech feature space through calculation or conversion mechanism, so that the emotional features derived from language information and the acoustic expression of speech generation task maintain structural compatibility. The cross-modal alignment operation can combine deep neural network mapping, contrast learning structure, feature transformation network or joint encoding mechanism to generate an acoustic feature basis vector, which is an intermediate expression with speech feature space expression ability, and retains the mapping structure of emotional information in the speech feature dimension.
[0121] The "thinking mode mechanism" is a modeling method that simulates the multi-modal information integration and dynamic adjustment strategy in human cognitive process, emphasizing information fusion and expression driven based on "context memory, time perception and target state constraint". This mechanism gives the model dynamic adaptive generation capability facing semantic target by constructing a cross-modal representation unified space, introducing an emotion-driven adjustment function, and constructing a time evolution channel based on self-attention. Compared with traditional static modeling methods, the thinking mode mechanism supports coupling modeling of language implicit emotional changes and speech time sequence expression, and automatically evaluating context weights at different time nodes to realize the unity of style continuity and expression tension.
[0122] Fusing the acoustic feature basis vector and the time state code based on the thinking mode mechanism refers to dynamically combining the static expression of speech features and the dynamic structure of time dimension by simulating the process of multi-modal information integration in human thinking activity. The time state code is usually based on the timestamp, rhythm control information or time position code generated by the speech sequence, and the fusion operation can be realized by position code network, dynamic adjustment mechanism or attention fusion network. The generated initial acoustic feature frame has time and content dual expression ability, providing a basis for subsequent continuous acoustic feature generation.
[0123] More specifically, the information fusion in the thinking mode mechanism not only includes the combination of static semantics and dynamic time, but also includes the multi-dimensional re-labeling processing of the "emotion-driven tensor", which is based on the cross-modal attention path, introduces a weak prediction of the future target state at the initial frame generation, and strengthens the trend expression ability of the speech frame in the dimensions of rhythm, tone, energy distribution, etc. The time state encoding establishes a coupling mapping with the emotion control quantity through the position perception network, ensuring that the generation of each frame is not only related to its position, but also consistent with the emotional evolution direction under the global context.
[0124] Starting from the initial acoustic feature frame, the thinking mode mechanism is iteratively executed based on the time sequence evolution operation, which simulates the continuous evolution process of speech in human language expression, dynamically generates subsequent acoustic feature expressions by combining the previous acoustic feature information and the current time state. The time sequence evolution operation can be realized through a sequence generation network, an autoregressive model, a recursive structure or a prediction enhancement module. The generated acoustic feature frame sequence is a sequence structure with complete time continuity, speech expression logic and emotional consistency, reflecting the acoustic profile and semantic emotion of the expected speech expression.
[0125] During the evolution process, the thinking mode mechanism constructs a dynamic decision-making path for "expression stability". Each frame generation is guided by "emotion consistency score" and "semantic trend prediction result", and the candidate feature generation path is selected through an attention-based scoring function. This mechanism can actively avoid sudden changes caused by emotional jumps, ensuring the audibility and naturalness of the audio sequence. In addition, historical context frames will dynamically participate in the construction of the current frame through a sliding window mechanism, avoiding expression breaks or emotional expression distortion.
[0126] Performing feature normalization operation on the acoustic feature frame sequence means to ensure expression accuracy while unifying the amplitude, scale, distribution or statistical properties of acoustic features. Common normalization operations include batch normalization, layer normalization, dynamic amplitude control or feature standardization techniques. The generated audio feature sequence has stable, continuous and natural acoustic expression ability, which meets the input requirements of the speech synthesis process.
[0127] In addition, the thinking mode mechanism can also maintain the consistency of the emotional label of the output feature in the normalization stage. By comparing the target feature distribution of the real emotional sample, the emotional bias trend of the generated feature is corrected. Especially in the mixed expression of multiple emotions or the ambiguous tone at the end of the sentence, the semantic emotional gravity center can be returned to normal, ensuring that the whole speech style is uniform, the rhythm is coordinated, and the sound quality is natural.
[0128] By the above steps, the emotion control information generated based on the text prompt can be converted into an audio feature sequence with complete emotion expression capability and time continuity through structure mapping, information fusion and time sequence modeling, breaking through the technical limitations of traditional speech synthesis methods such as disconnection between emotion expression and speech generation, lack of dynamic evolution structure, and unstable distribution of acoustic features, improving the naturalness, coherence and emotion authenticity of speech generation, and meeting the demand for high-quality speech generation in various emotions and complex contexts.
[0129] In one embodiment, the above step S50 comprises:
[0130] S501, determining the duration feature of each phoneme in the phoneme sequence to generate phoneme duration information;
[0131] S502, determining the total duration of the phoneme sequence and the total duration of the audio feature sequence;
[0132] S503, when the total duration of the audio feature sequence is less than the total duration of the phoneme sequence, performing a linear interpolation frame expansion operation on the audio feature sequence to generate an adjusted audio feature sequence, and dividing the adjusted audio feature sequence into an audio feature frame segment sequence based on the phoneme duration information;
[0133] S504, when the total duration of the phoneme sequence is less than the total duration of the audio feature sequence, performing an equal-proportion duration expansion operation on the phoneme sequence to generate an adjusted phoneme sequence, and dividing the audio feature sequence into an audio feature frame segment sequence based on the duration feature of the adjusted phoneme sequence;
[0134] S505, establishing a mapping relationship between the phoneme sequence unit and the audio feature frame segment sequence to generate a time mapping table;
[0135] S506, binding each phoneme sequence unit with the corresponding audio feature frame segment to generate a binding unit;
[0136] S507, adding a timestamp label to each binding unit and reorganizing all binding units in chronological order to generate a time-aligned sequence.
[0137] In this embodiment, the phoneme sequence is a discrete symbol sequence obtained by mapping text content to express the structure of speech content, where each phoneme corresponds to a language pronunciation unit, usually based on the International Phonetic Alphabet, a speech annotation system or a custom phoneme set. The audio feature sequence is a continuous acoustic feature sequence generated dynamically in combination with emotion control information and time, which is used to reflect the timing, spectrum and voice quality parameters of speech. There are independent generation differences in the time dimension between the two, which need to be unified through time alignment operation.
[0138] The duration feature of each phoneme in the phoneme sequence is determined, which refers to obtaining the time occupancy information of each phoneme in the expected speech expression based on language rules, pronunciation prediction model or data-driven method. The duration feature can be estimated by statistical learning, speech sample analysis or acoustic model to form phoneme duration information. The phoneme duration information usually records the duration data of each phoneme in units of time slices, frame numbers or milliseconds.
[0139] The total duration of the phoneme sequence and the total duration of the audio feature sequence are determined, which refers to calculating the cumulative total duration of the phoneme duration information and the actual covered time of the audio feature sequence respectively, and evaluating the time consistency by comparing the two. The total duration of the audio feature sequence can be calculated according to the number of feature frames and the time interval of a single frame.
[0140] When the total duration of the audio feature sequence is less than the total duration of the phoneme sequence, linear interpolation frame expansion operation is performed on the audio feature sequence, which refers to inserting additional frames in the time dimension of the original audio feature sequence by interpolation technology to maintain the smooth transition of acoustic features and generate an adjusted audio feature sequence. Linear interpolation frame expansion can generate new feature points based on adjacent feature frame data according to the time distribution rule. The expanded sequence meets the time length requirement. Based on the phoneme duration information, the adjusted audio feature sequence is segmented into an audio feature frame segment sequence, which refers to cutting the expanded audio feature sequence into a continuous frame segment set corresponding to the phoneme sequence structure according to the time boundary of the phoneme duration information, forming an audio feature frame segment sequence.
[0141] When the total duration of the phoneme sequence is less than the total duration of the audio feature sequence, an equal-proportion duration expansion operation is performed on the phoneme sequence, which refers to enlarging the phoneme duration as a whole while keeping the time proportion of each phoneme unchanged, so that the total duration of the phoneme sequence is consistent with the total duration of the audio feature sequence. The expansion method can be realized by time proportion transformation, duration interpolation or speech rhythm reconstruction. Based on the duration feature of the adjusted phoneme sequence, the audio feature sequence is segmented into an audio feature frame segment sequence, which has the same process as before, ensuring the structural mapping relationship between phonemes and audio features.
[0142] A mapping relationship between the phoneme sequence unit and the audio feature frame segment sequence is established, which refers to one-to-one correspondence between each phoneme sequence unit, i.e. the basic pronunciation unit with phoneme symbol and duration attribute, and the audio feature frame segment in the corresponding time period, forming a time mapping table. The time mapping table records the time, structure and content association information between phonemes and audio features.
[0143] Each phoneme sequence unit is bound to the corresponding audio feature frame segment to generate a binding unit, which refers to encapsulating the phoneme unit and the audio feature frame segment it covers into an integrated expression structure based on the time mapping table. The binding unit has the joint expression capability of phoneme information, time information and acoustic feature information.
[0144] A timestamp mark is added to each binding unit, and all binding units are reorganized in time sequence to generate a time-aligned sequence, which is a start-end time label attached to the binding unit to ensure the consistency of the overall time sequence. The reorganization operation is based on timestamp sorting to form a time-aligned sequence with clear structure and complete timing. The time-aligned sequence can be used in subsequent speech synthesis, decoding or quality control steps.
[0145] The embodiment can effectively solve the problems of inconsistent time length and unclear structure correspondence between the phoneme sequence and the audio feature sequence generated independently, ensure the time synchronization and structure unity of language content and acoustic features in the speech expression process, improve the naturalness, accuracy and coherence of speech synthesis, avoid the situation of rhythm misplacement, incoherent pronunciation or lack of emotional expression in speech expression, and meet the demand for high-quality and emotionally rich speech generation.
[0146] In one embodiment, the above step S60 includes:
[0147] S601, extracting an audio feature frame segment from each binding unit of the time-aligned sequence;
[0148] S602, performing a time sequence reorganization operation on the audio feature frame segment to generate speech decoder input features;
[0149] S603, generating original waveform data by iteratively processing the speech decoder input features through an autoregressive waveform generation layer of a speech decoder;
[0150] S604, performing dynamic range compression and clipping processing on the original waveform data to generate speech waveform.
[0151] In the embodiment, the time-aligned sequence is an ordered data set obtained by structurally binding phoneme information and corresponding audio feature frame segments, and has three types of information: phoneme symbols, timestamps and audio features. The time-aligned sequence is input into a speech decoder, which is a neural network module specially used for speech signal recovery and reconstruction. The speech decoder converts discrete audio features into continuous speech signals through a learned generation mechanism.
[0152] Extracting an audio feature frame segment from each binding unit of the time-aligned sequence means traversing the binding units in the time-aligned sequence one by one to extract the audio feature frame segment data contained therein. The audio feature frame segment is a continuous acoustic parameter sequence, usually including mel-frequency spectrum, cepstrum coefficient, amplitude information, etc. The extraction operation preserves the time sequence and frame structure to ensure the integrity of the feature data.
[0153] The time sequence reorganization operation is performed on the audio feature frame segment to generate speech decoder input features. The extracted audio feature frame segments are sequentially spliced into a complete time sequence according to the timestamp markers. The time sequence reorganization ensures that the order of frame segments corresponding to different phonemes is consistent with the original language structure. The generated speech decoder input features are continuous feature sets for speech synthesis tasks, with the characteristics of time sequence coherence, data integrity, and semantic correspondence.
[0154] The speech decoder input features are iteratively processed by an autoregressive waveform generation layer of the speech decoder to generate original waveform data. The autoregressive waveform generation layer is a recursive processing module based on a neural network architecture, which typically includes residual convolution, gated recurrent unit, or attention mechanism. Through frame-by-frame prediction, the waveform data at the next time point is dynamically generated based on the historically generated waveform fragments and the current input features. The entire process is recursively unfolded in the time dimension until the complete original waveform data is generated. The original waveform data is a continuous digital signal that has not been filtered or compressed, retaining the basic tone and timing characteristics of the speech.
[0155] Dynamic range compression and clipping processing are performed on the original waveform data to generate speech waveforms. Dynamic range compression reduces the amplitude difference between strong and weak signals in the speech, improving the balance and clarity of the overall audio. Clipping processing limits the maximum amplitude of the signal to avoid signal distortion, clipping, or overload. The combination of the two can optimize speech quality, improve listening comfort, and enhance device compatibility. The generated speech waveform is the final audio signal that meets the actual application requirements and can be directly used for speech output, storage, or transmission.
[0156] The above steps can effectively improve the accuracy and naturalness of the conversion of audio features to actual speech signals in the speech synthesis system, avoiding unclear pronunciation, speech rhythm disorder, or tone distortion caused by the lack of structural constraints between audio feature sequences and waveform signals. Dynamic range compression and clipping processing further enhance the expression effect of the speech waveform, ensuring that the output speech has high clarity, accurate emotional expression, natural listening, and is suitable for playing on various devices, meeting the actual application requirements of high-quality speech synthesis systems.
[0157] In one embodiment, after the above step S60, the method further comprises:
[0158] S701, inputting the speech waveform, text prompt, and source text into a multi-modal analysis module, performing an emotion recognition task through the multi-modal analysis module, and generating an emotion consistency score;
[0159] S702, when the emotion consistency score is lower than a preset score threshold, adjusting the emotion control vector to generate an adjusted emotion control vector;
[0160] S703, generate an adjusted audio feature sequence based on the adjusted emotion control vector;
[0161] S704, re-execute the time alignment operation based on the adjusted audio feature sequence to generate a new time alignment sequence;
[0162] S705, input the new time alignment sequence into the speech decoder to generate an optimized speech waveform.
[0163] In this embodiment, the time alignment sequence refers to a structured data sequence generated based on the binding of phoneme information and audio feature frame segments and the completion of timestamp marking. After inputting the speech decoder and generating a speech waveform, the obtained speech waveform is an initial speech signal synthesized based on the input text prompt and the source text content.
[0164] Inputting the speech waveform, the text prompt, and the source text into the multi-modal analysis module refers to jointly inputting three data of different sources but semantically related into an analysis system with cross-modal feature fusion capability. The speech waveform is an audio signal to be evaluated, the text prompt is a user-specified emotional expression intention, and the source text is the original text information corresponding to the speech content. The multi-modal analysis module can cooperatively reason based on audio and text data forms, fuse and extract emotional features and semantic features, and realize a more comprehensive emotion recognition task.
[0165] Through the multi-modal analysis module, an emotion consistency score is generated, which refers to the analysis module combining the speech signal features in the speech waveform and the semantic expression in the text information to determine whether the generated speech accurately conveys the emotional type and expression intensity specified in the text prompt. The emotion consistency score reflects the matching degree of the speech content and the emotion instruction. The higher the score, the more the speech expression conforms to the expectation, and the lower the score, the more the deviation.
[0166] When the emotion consistency score is lower than a preset score threshold, the emotion control vector is adjusted to generate an adjusted emotion control vector. The preset score threshold is the lowest acceptable standard defined by the system. A score lower than this standard indicates that the speech emotion expression has a deviation, which needs to be dynamically corrected by adjusting the emotion control vector. The adjusted emotion control vector can be generated based on the emotion deviation type, the deviation degree, or the user-set strategy to ensure higher expression accuracy in the next synthesis process.
[0167] Based on the adjusted emotion control vector, an adjusted audio feature sequence is generated, which refers to inputting the updated emotion control vector into the thinking mode mechanism to generate a new audio feature sequence through cross-modal feature fusion and time series modeling. The adjusted audio feature sequence has differences in length, intensity, tone, etc. compared to the initial audio feature sequence, so as to better match the target emotional expression needs.
[0168] Based on the adjusted audio feature sequence, the time alignment operation is re-executed to generate a new time alignment sequence, which refers to the time sequence matching, frame segment segmentation, time mapping, unit binding, and timestamp marking operations of the audio feature sequence and the phoneme sequence, to generate a structured new time alignment sequence, ensuring that the adjusted audio feature sequence and the semantic structure of the synthesized speech are consistent.
[0169] The new time alignment sequence is input into the speech decoder to generate an adjusted speech waveform, which refers to re-performing speech decoding operations based on updated time alignment data, and the generated adjusted speech waveform is more consistent with the target expression requirements specified by the text prompt in terms of emotional expression, naturalness of voice quality, and speech rhythm, etc., improving the quality of the final speech output.
[0170] Example: In a medical health information interaction system, for patient psychological intervention, health education or remote consultation process, combined with voice intelligent generation technology, doctors can input structured source text content and text prompt information in natural language form, and the system automatically synthesizes speech output with specified emotional expression based on the input content, improves the emotional transmission effect of doctor-patient communication, relieves the anxiety of patients, and optimizes the effect of health information transmission.
[0171] In the specific operation process, first, the doctor inputs the source text content through the terminal interface, such as "Please keep optimistic, cooperate with treatment, we will work together to improve your condition", and inputs the text prompt information, indicating the system to generate the corresponding speech with "mild, encouraging" emotional expression.
[0172] After the system receives the source text and text prompt information, it processes the text prompt using the language model, disassembles the text prompt into a word sequence through word segmentation technology, and further converts it into a word embedding vector sequence, which is input into the encoding layer of the language model to extract hidden state sequence features. After global pooling operation, a sentiment control vector is generated, which quantitatively expresses the emotional instructions "mild, encouraging" in the text prompt.
[0173] Subsequently, the system processes the source text based on the thought chain mechanism. First, the source text content is segmented into an ordered word unit sequence, and the word unit sequence is gradually converted into an initial phoneme symbol sequence using phoneme mapping rules. Through the thought chain mechanism, the initial phoneme symbol sequence is executed in a hierarchical progressive expansion operation to enrich the language structure of speech expression. Combined with the phoneme boundary optimization strategy, a standard phoneme sequence is generated to ensure the clear and accurate phonetic structure of the speech content.
[0174] The system further processes the emotion control vector based on the thinking mode mechanism, first performs a cross-modal alignment operation to map the emotional information to the acoustic feature space, generates an acoustic feature basis vector, fuses the time state encoding information, and generates an initial acoustic feature frame. Taking this as a starting point, the system iteratively constructs a time evolution process through the thinking mode mechanism to form a coherent sequence of acoustic feature frames. Finally, through feature normalization operations, the system outputs an audio feature sequence that meets the specified emotional characteristics.
[0175] After completing the construction of the phoneme sequence and the audio feature sequence, the system performs a time alignment operation to calculate the duration information of each phoneme in the phoneme sequence and obtain the total duration difference between the phoneme sequence and the audio feature sequence. If the audio feature sequence is shorter, the system uses linear interpolation frame expansion technology to expand the audio feature sequence and performs frame segmentation based on the phoneme duration information. If the phoneme sequence is shorter, the system uses proportional length expansion operations to correct the phoneme sequence and performs frame segmentation accordingly. The system then establishes a mapping relationship between the phoneme sequence units and the audio feature frame segment sequence, generates a time mapping table, and further binds the phoneme sequence units with the corresponding audio feature frames and adds timestamp information to reorganize and generate a time-aligned sequence in chronological order.
[0176] The system inputs the time-aligned sequence into the speech decoder, extracts the audio feature frame segments in each binding unit, performs time sequence reorganization to generate speech decoder input features, iteratively generates original waveform data based on the autoregressive waveform generation layer, completes dynamic range compression and clipping processing, and outputs speech waveforms with "mild and encouraging" emotions as the output content of the doctor's speech information.
[0177] To ensure the accuracy of the speech emotion expression, the system simultaneously inputs the speech waveform, the text prompt, and the source text into the multi-modal analysis module, jointly analyzes the audio and text information, performs an emotion recognition task, and generates an emotion consistency score. If the score is lower than the system's set threshold, it indicates that the speech expression deviates from the expectation. The system automatically adjusts the emotion control vector, updates the generated audio feature sequence, re-performs the time alignment operation, generates a new time-aligned sequence, inputs it into the speech decoder, iteratively optimizes the adjusted speech waveform, and ensures that the speech expression accurately reflects the doctor's intention, thereby improving the effectiveness of the medical and health information communication and the psychological acceptance of patients.
[0178] In the field of financial technology business, for scenarios such as intelligent customer service, financial product promotion, risk warning, or user emotion soothing, combined with natural language-driven personalized voice generation technology, by inputting structured source text information and text prompt content, the system can generate financial voice output with free expression style and delicate emotion control according to actual business needs, significantly improving user communication experience and the emotional accuracy of information transmission, alleviating customer anxiety caused by information asymmetry or single service, and enhancing business trust.
[0179] In a specific implementation process, the financial service system receives source text information such as "According to your risk preference, it is recommended to pay attention to low-volatility financial products in this quarter", and inputs text prompt information indicating the emotional requirements of "smooth, professional, and trustworthy". The system automatically starts the language understanding and emotional expression collaborative processing flow.
[0180] Firstly, the system performs word segmentation operation on the text prompt information based on the language model, disassembles the text prompt into word sequence, and further converts it into word embedding vector sequence. The system inputs the encoding layer of the language model, extracts hidden state sequence features, generates emotion control vector through global pooling operation, and the vector explicitly expresses the emotional instructions of "smooth, professional, and trustworthy". This ensures that the system accurately reflects the specified emotional style of financial services in the subsequent speech generation process.
[0181] For the source text information, the system performs progressive processing based on the thought chain mechanism. Firstly, the source text content is divided into an ordered word unit sequence. The initial phoneme symbol sequence is constructed using the phoneme mapping relationship. The speech structure information is enriched through multi-level expansion operation. Finally, the standard phoneme sequence with clear semantic expression and complete speech structure is generated through phoneme boundary optimization process, which guarantees the accuracy and naturalness of the generated speech in the expression of professional terms and financial terms.
[0182] The system further processes the emotion control vector through the thought mode mechanism. Firstly, the cross-modal alignment operation is performed to map the abstract emotional information to the acoustic feature space to form the acoustic feature basis vector. The initial acoustic feature frame is generated by fusing the time state encoding information. Based on the feature frame, the system iteratively performs the time evolution process to dynamically construct the acoustic feature frame sequence. Through feature normalization operation, the audio feature sequence expressing the "smooth, professional, and trustworthy" style in the context of financial services is output.
[0183] After completing the construction of the phoneme sequence and the audio feature sequence, the system performs time alignment operation according to the phoneme duration and the total duration of the audio feature. If the audio feature sequence is short, the system extends the audio feature sequence through linear interpolation frame expansion technology, and divides it into audio feature frame segment sequence according to the phoneme duration information. If the phoneme sequence is short, the system adjusts the phoneme sequence through equal proportion time length expansion, synchronously divides the audio feature frame segment sequence, establishes the mapping relationship between the phoneme sequence unit and the audio feature frame segment, generates the time mapping table, and recombines all binding units to form the time alignment sequence combined with the time stamp information.
[0184] The system inputs the time-aligned sequence into the speech decoder, extracts the audio feature frame segment in the binding unit, performs a time sequence reorganization operation, generates speech decoder input features, generates original waveform data frame by frame based on an autoregressive waveform generation layer, further performs dynamic range compression and clipping processing, and outputs speech waveforms, finally generating financial speech output embodying the "stable, professional, and trustworthy" emotional style.
[0185] To ensure the consistency of emotional expression in financial information transmission with user expectations, the system synchronously inputs speech waveforms, text prompts, and source text into a multi-modal analysis module, combines audio, text, and context information, performs an emotional recognition task, generates an emotional consistency score, and if the score is lower than the system's set trust threshold, the system automatically adjusts the emotional control vector, regenerates the audio feature sequence, performs a new time alignment process, inputs the speech decoder, and generates optimized speech waveforms, ensuring that the financial speech output accurately transmits business information while having delicate and trustworthy emotional expression effects, improving user service experience and the professionalism and affinity of financial product communication.
[0186] The above steps realize dynamic feedback and adaptive optimization of speech generation quality, overcome the problem that emotional expression deviation in a single speech generation process cannot be corrected in time, the multi-modal analysis module uses audio and text information to jointly evaluate the emotional accuracy of generated speech, ensures that the output speech meets the emotional instruction requirements set by the user, and the adjustment process optimizes the emotional control vector and the audio feature sequence, combines the re-executed time alignment and speech decoding operations, effectively improves the expression flexibility and accuracy of the generated effect of the speech synthesis system, enhances the consistency of speech content and target emotion, and improves the human-computer interaction experience and practical value of the speech synthesis system.
[0187] In an embodiment, a thought chain and thought modality assisted speech generation device is provided, which corresponds to the thought chain and thought modality assisted speech generation method in the above embodiment. Referring to Figure 3 , Figure 3 A functional module schematic diagram of a preferred embodiment of the thought chain and thought modality assisted speech generation device of the present application. The input analysis module 10, the emotional modeling module 20, the phoneme generation module 30, the audio feature generation module 40, the time alignment module 50, and the speech synthesis module 60. The detailed description of each functional module is as follows:
[0188] The input analysis module 10 is used to receive source text and text prompts for specifying emotional expression;
[0189] The emotional modeling module 20 is used to input the text prompts into the language model and generate an emotional control vector through the language model;
[0190] The phoneme generation module 30 is configured to generate a phoneme sequence based on the thought chain mechanism processing the source text.
[0191] The audio feature generation module 40 is configured to generate an audio feature sequence based on the thought mode mechanism processing the emotion control vector.
[0192] The time alignment module 50 is configured to perform a time alignment operation on the phoneme sequence and the audio feature sequence to generate a time alignment sequence.
[0193] The speech synthesis module 60 is configured to input the time alignment sequence into a speech decoder to generate a speech waveform.
[0194] In an embodiment, the emotion modeling module 20 is specifically configured to:
[0195] perform a word segmentation operation on the text prompt to generate a word sequence;
[0196] convert the word sequence into a word embedding vector sequence;
[0197] input the word embedding vector sequence into a language model to generate a hidden state sequence through an encoding layer of the language model;
[0198] perform a global pooling operation on the hidden state sequence to generate an emotion control vector.
[0199] In an embodiment, the phoneme generation module 30 is specifically configured to:
[0200] divide the source text into an ordered word unit sequence;
[0201] perform a phoneme symbol mapping operation on the ordered word unit sequence to generate an initial phoneme symbol sequence;
[0202] analyze semantic features of the phonemes in a context based on an ordered word unit corresponding to each phoneme in the initial phoneme symbol sequence;
[0203] based on phoneme coordination pronunciation strategies between the semantic features and adjacent phonemes, locate positions where a connecting phoneme needs to be inserted, and insert the connecting phoneme at the corresponding positions to generate a sequence after insertion of the connecting phoneme;
[0204] based on a result of emotion state classification or a tone turning point of a context semantic of the ordered word unit corresponding to the phoneme, insert a prosodic emphasis marker at a position where a prosodic emphasis probability is greater than a preset threshold or a syntactic boundary, to generate a sequence after insertion of the prosodic marker;
[0205] inserting the prosodic mark into the post-sequence as a new current phoneme sequence, iteratively performing the semantic feature analysis operation, the phoneme insertion operation and the prosodic emphasis mark insertion operation, each time taking the phoneme sequence generated in the last iteration as input, updating the context information and continuing processing until covering all ordered word units corresponding to the phonemes in the initial phoneme symbol sequence, and generating an extended phoneme sequence;
[0206] performing a phoneme boundary optimization operation on the extended phoneme sequence to generate a phoneme sequence.
[0207] In an embodiment, the audio feature generation module 40 is specifically configured to:
[0208] performing a cross-modal alignment operation on the emotion control vector to generate an acoustic feature base vector;
[0209] fusing the acoustic feature base vector with the time state encoding based on the thinking modal mechanism to generate an initial acoustic feature frame;
[0210] taking the initial acoustic feature frame as a starting point, iteratively performing a time evolution operation based on the thinking modal mechanism to generate an acoustic feature frame sequence;
[0211] performing feature normalization processing on the acoustic feature frame sequence to generate an audio feature sequence.
[0212] In an embodiment, the time alignment module 50 is specifically configured to:
[0213] determining the duration feature of each phoneme in the phoneme sequence to generate phoneme duration information;
[0214] determining the total duration of the phoneme sequence and the total duration of the audio feature sequence;
[0215] when the total duration of the audio feature sequence is less than the total duration of the phoneme sequence, performing a linear interpolation frame expansion operation on the audio feature sequence to generate an adjusted audio feature sequence, and dividing the adjusted audio feature sequence into an audio feature frame segment sequence based on the phoneme duration information;
[0216] when the total duration of the phoneme sequence is less than the total duration of the audio feature sequence, performing an equal-proportion duration expansion operation on the phoneme sequence to generate an adjusted phoneme sequence, and dividing the audio feature sequence into an audio feature frame segment sequence based on the duration feature of the adjusted phoneme sequence;
[0217] establishing a mapping relationship between the phoneme sequence units and the audio feature frame segment sequence to generate a time mapping table;
[0218] binding each phoneme sequence unit with the corresponding audio feature frame segment to generate a binding unit;
[0219] Add a time stamp mark to each binding unit, and reorganize all the binding units in time sequence to generate a time-aligned sequence.
[0220] In an embodiment, the speech synthesis module 60 is specifically configured to:
[0221] Extract an audio feature frame segment from each binding unit of the time-aligned sequence;
[0222] Perform a time series reorganization operation on the audio feature frame segment to generate speech decoder input features;
[0223] Iteratively process the speech decoder input features through an autoregressive waveform generation layer of a speech decoder to generate raw waveform data;
[0224] Perform dynamic range compression and clipping processing on the raw waveform data to generate a speech waveform.
[0225] In an embodiment, the speech synthesis module 60 is specifically configured to:
[0226] Input the speech waveform, the text prompt, and the source text into a multi-modal analysis module, perform a sentiment recognition task through the multi-modal analysis module, and generate a sentiment consistency score;
[0227] When the sentiment consistency score is lower than a preset score threshold, adjust the emotion control vector to generate an adjusted emotion control vector;
[0228] Generate an adjusted audio feature sequence based on the adjusted emotion control vector;
[0229] Re-execute a time alignment operation based on the adjusted audio feature sequence to generate a new time-aligned sequence;
[0230] Input the new time-aligned sequence into the speech decoder to generate an optimized speech waveform.
[0231] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. The processor of the computer device is configured to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external user terminal through a network connection. The computer program is executed by the processor to implement the functions or steps of the server side of the thought chain and thought modal auxiliary speech generation method.
[0232] In one embodiment, a computer device is provided, which can be a user terminal, and its internal structure diagram can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with an external server through a network connection. The computer program is executed by the processor to implement the functions or steps of a thought chain and thought mode auxiliary voice generation method on the user terminal side.
[0233] In one embodiment, a computer device is provided, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the following steps:
[0234] Obtain input data and encode the input data to generate an initial data representation;
[0235] Perform semantic analysis on the input data to generate semantic features;
[0236] Input the semantic features into an attribute encoder to generate a dynamic attribute representation vector;
[0237] Fuse the dynamic attribute representation vector with the initial data representation to obtain fused features;
[0238] Generate intermediate spectral representations based on the fused features through a decoder;
[0239] Input the intermediate spectral representations into a vocoder to convert the intermediate spectral representations into target signals through the vocoder.
[0240] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0241] Obtain input data and encode the input data to generate an initial data representation;
[0242] Perform semantic analysis on the input data to generate semantic features;
[0243] Input the semantic features into an attribute encoder to generate a dynamic attribute representation vector;
[0244] Fusing the dynamic attribute characterization vector with the initial data characterization to obtain a fused feature;
[0245] Generating, by a decoder, an intermediate spectral characterization based on the fused feature;
[0246] Inputting the intermediate spectral characterization into a vocoder, and converting, by the vocoder, the intermediate spectral characterization into a target signal.
[0247] It should be noted that the functions or steps described above with respect to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0248] Those skilled in the art can understand that all or part of the processes in the foregoing method embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the foregoing embodiments of the method can be included. In each embodiment provided in the present application, any reference to a memory, storage, database or other medium can include a non-volatile and / or volatile memory. The non-volatile memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM) or a flash memory. The volatile memory can include a random access memory (RAM) or an external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0249] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified. In actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0250] It should be explained that if the software tools or components of other companies appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for generating an auxiliary voice for a thought chain and a thought mode, characterized by, The method comprises the following steps: receiving a source text and a text prompt for specifying emotional expression; inputting the text prompt into a language model to generate an emotion control vector through the language model; processing the source text based on a thinking chain mechanism to generate a phoneme sequence, comprising: segmenting the source text into an ordered word unit sequence; performing a phoneme symbol mapping operation on the ordered word unit sequence to generate an initial phoneme symbol sequence; analyzing the semantic features of the phonemes in the context based on the ordered word unit corresponding to each phoneme in the initial phoneme symbol sequence; based on the semantic features and the phoneme coarticulation strategy between adjacent phonemes, locating the position where a connecting phoneme needs to be inserted, and inserting the connecting phoneme at the corresponding position to generate a sequence after inserting the connecting phoneme; based on the emotion state classification result or the intonation turning point of the context semantics of the phoneme corresponding ordered word unit, inserting a prosodic emphasis marker at a position where the stress probability is greater than a preset threshold or a syntactic boundary to generate a sequence after inserting the prosodic marker; taking the sequence after inserting the prosodic marker as a new current phoneme sequence, iteratively performing semantic feature analysis, connecting phoneme insertion and prosodic emphasis marker insertion, each iteration taking the phoneme sequence generated in the last iteration as input, updating the context information and continuing processing until all ordered word units corresponding to the phonemes in the initial phoneme symbol sequence are covered, to generate an extended phoneme sequence; performing a phoneme boundary optimization operation on the extended phoneme sequence to generate a phoneme sequence; processing the emotion control vector based on a thinking mode mechanism to generate an audio feature sequence, comprising: performing a cross-modal alignment operation on the emotion control vector to generate an acoustic feature base vector; based on the thinking mode mechanism, fusing the acoustic feature base vector and the time state encoding to generate an initial acoustic feature frame; taking the initial acoustic feature frame as a starting point, iteratively performing a time evolution operation based on the thinking mode mechanism to generate an acoustic feature frame sequence; performing feature normalization processing on the acoustic feature frame sequence to generate an audio feature sequence; performing a time alignment operation on the phoneme sequence and the audio feature sequence to generate a time-aligned sequence; inputting the time-aligned sequence into a speech decoder to generate a speech waveform.
2. The method of claim 1, wherein the method further comprises: inputting the text prompt into a language model to generate an emotion control vector through the language model, comprising: performing word segmentation processing on the text prompt to generate a word sequence; converting the word sequence into a word embedding vector sequence; inputting the word embedding vector sequence into the language model to generate a hidden state sequence through the encoding layer of the language model; performing a global pooling operation on the hidden state sequence to generate an emotion control vector.
3. The method of claim 1, wherein the method further comprises: performing a time alignment operation on the phoneme sequence and the audio feature sequence to generate a time-aligned sequence, comprising: determining the duration characteristics of each phoneme in the phoneme sequence to generate phoneme duration information; determining the total duration of the phoneme sequence and the total duration of the audio feature sequence; When the total time length of the audio feature sequence is less than the total time length of the phoneme sequence, performing a linear interpolation frame expansion operation on the audio feature sequence to generate an adjusted audio feature sequence, and segmenting the adjusted audio feature sequence into an audio feature frame segment sequence based on the phoneme time length information; When the total time length of the phoneme sequence is less than the total time length of the audio feature sequence, performing an equal-proportion time length expansion operation on the phoneme sequence to generate an adjusted phoneme sequence, and segmenting the audio feature sequence into an audio feature frame segment sequence based on the duration feature of the adjusted phoneme sequence; establishing a mapping relationship between the phoneme sequence units and the audio feature frame segment sequence to generate a time mapping table; binding each phoneme sequence unit with the corresponding audio feature frame segment to generate a binding unit; adding a timestamp label to each binding unit and reorganizing all binding units in chronological order to generate a time-aligned sequence.
4. The method of claim 1, wherein the method further comprises: inputting the time-aligned sequence into a speech decoder to generate a speech waveform, including: extracting an audio feature frame segment from each binding unit of the time-aligned sequence; performing a time series reorganization operation on the audio feature frame segment to generate a speech decoder input feature; iteratively processing the speech decoder input feature through an autoregressive waveform generation layer of the speech decoder to generate raw waveform data; performing dynamic range compression and clipping processing on the raw waveform data to generate a speech waveform.
5. The method of generating voice for thought chain and thought modality as claimed in claim 1, wherein, After inputting the time-aligned sequence into a speech decoder to generate a speech waveform, further comprising: inputting the speech waveform, the text prompt, and the source text into a multi-modal analysis module to perform a sentiment recognition task through the multi-modal analysis module to generate a sentiment consistency score; when the sentiment consistency score is lower than a preset score threshold, adjusting the sentiment control vector to generate an adjusted sentiment control vector; generating an adjusted audio feature sequence based on the adjusted sentiment control vector; re-performing the time alignment operation based on the adjusted audio feature sequence to generate a new time-aligned sequence; inputting the new time-aligned sequence into the speech decoder to generate an optimized speech waveform.
6. A thought chain and thought modality assisted speech generation device, characterized by, The thought chain and thought modality auxiliary speech generation device comprises: an input analysis module for receiving a source text and a text prompt for specifying emotional expression; an emotion modeling module for inputting the text prompt into a language model to generate a sentiment control vector through the language model; The phoneme generation module is configured to generate a phoneme sequence based on a thought chain mechanism, including: segmenting the source text into an ordered word unit sequence; performing a phoneme symbol mapping operation on the ordered word unit sequence to generate an initial phoneme symbol sequence; analyzing semantic features of the phonemes in a context based on an ordered word unit corresponding to each phoneme in the initial phoneme symbol sequence; positioning a position where a connecting phoneme needs to be inserted based on the semantic features and a phoneme coarticulation strategy between adjacent phonemes, and inserting the connecting phoneme at the corresponding position to generate a sequence after insertion of the connecting phoneme; inserting a prosodic emphasis marker at a position where a prosodic probability is greater than a preset threshold or a syntactic boundary based on a sentiment state classification result or a tone turning point of a context semantic of the ordered word unit corresponding to the phoneme to generate a sequence after insertion of the prosodic marker; taking the sequence after insertion of the prosodic marker as a new current phoneme sequence, iteratively performing the semantic feature analysis operation, the connecting phoneme insertion operation, and the prosodic emphasis marker insertion operation, taking the phoneme sequence generated in the last iteration as an input, and updating context information for continuous processing until all ordered word units corresponding to the phonemes in the initial phoneme symbol sequence are covered to generate an extended phoneme sequence; and performing a phoneme boundary optimization operation on the extended phoneme sequence to generate a phoneme sequence. The audio feature generation module is configured to generate an audio feature sequence based on a thought modality mechanism by processing the emotion control vector, including: performing a cross-modal alignment operation on the emotion control vector to generate an acoustic feature base vector; fusing the acoustic feature base vector and a time state code based on the thought modality mechanism to generate an initial acoustic feature frame; taking the initial acoustic feature frame as a starting point, iteratively performing a time evolution operation based on the thought modality mechanism to generate an acoustic feature frame sequence; and performing feature normalization processing on the acoustic feature frame sequence to generate an audio feature sequence. The time alignment module is configured to perform a time alignment operation on the phoneme sequence and the audio feature sequence to generate a time-aligned sequence. The speech synthesis module is configured to input the time-aligned sequence into a speech decoder to generate a speech waveform.
7. A computer device, comprising: The computer device includes a memory, a processor, and a thought chain and thought modality assisted speech generation program stored on the memory and executable on the processor, and the thought chain and thought modality assisted speech generation program, when executed by the processor, implements the steps of the thought chain and thought modality assisted speech generation method of any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The storage medium stores a thought chain and thought modality assisted speech generation program, and the thought chain and thought modality assisted speech generation program, when executed by the processor, implements the steps of the thought chain and thought modality assisted speech generation method of any one of claims 1-5.
Citation Information
Patent Citations
Audio generation model training method and device, computer equipment and storage medium
CN119132271A
Adaptive traffic domain service voice generation method and system based on thinking chain fine-tuning large model
CN120126484A