Speech generation method based on multi-modal input and related device
By using a multimodal data manager and a context and control layer, the problem of insufficient utilization of multimodal data in existing speech synthesis technologies is solved, and high-quality, personalized speech generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 深圳市金政软件技术有限公司
- Filing Date
- 2025-09-17
- Publication Date
- 2026-05-05
AI Technical Summary
Existing speech synthesis technologies, when utilizing multimodal data, employ simple synthesis methods, making it difficult to meet the diverse needs of users.
By receiving multimodal input data and auxiliary input data, the multimodal data manager performs fusion processing, combines the correlation features of the auxiliary input data to generate multimodal fusion features, and processes them through the context and control layer and the speech generation layer to finally output high-quality speech synthesis.
It achieves in-depth utilization of multimodal data, and the generated speech is more in line with user needs. The speech generation strategy is flexible and diverse, meeting the requirements of personalization and coherence.
Smart Images

Figure CN121053996B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech generation method and related equipment based on multimodal input. Background Technology
[0002] With the rapid development of artificial intelligence, existing text-to-speech (TTS) technology has made significant progress. Existing TTS technology can now convert text into natural and fluent speech, and has been widely used in voice assistants, intelligent customer service, automatic translation and other fields.
[0003] To achieve diversity in speech synthesis, existing technologies have begun to use multimodal data to synthesize speech. However, the utilization of multimodal data in existing speech synthesis technologies is mostly limited to converting multimodal data into text to directly drive speech conversion or simply concatenating text before driving speech conversion. Existing speech synthesis technologies are simple in their speech synthesis methods and do not make sufficient use of multimodal data, making it difficult to meet user needs.
[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0005] This invention provides a speech generation method and related equipment based on multimodal input. The main objective of this invention is to solve the technical problems mentioned in the background section of the prior art.
[0006] The first aspect of this invention provides a speech generation method based on multimodal input, comprising:
[0007] The system receives multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data.
[0008] The multimodal input data and the auxiliary input data are input into a preset multimodal data manager. The multimodal input data is fused by the multimodal data manager, and the multimodal fusion features are obtained by combining the correlation features in the auxiliary input data.
[0009] The multimodal fusion features are input into the context and control layer. The multimodal fusion features are processed by the context state modeler, speech generation logic component and speech command generator in the context and control layer to obtain the target speech command. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller and a speech text generator.
[0010] The target speech command is input to the speech generation layer, and the speech parameter controller, speech synthesis scheduling component and speech stream controller of the speech generation layer process the target speech command to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine and a post-processor.
[0011] The target synthesized speech is output through an output layer, which includes at least one of the following: local output, online download and playback, transcoding and embedding into other multimedia files, pushing to a third-party system, and asynchronous processing.
[0012] In an optional embodiment of the first aspect of the present invention, the step of inputting the multimodal fusion features into the context and control layer, and processing the multimodal fusion features through the context state modeler, speech generation logic component, and speech command generator in the context and control layer to obtain the target speech command includes:
[0013] The multimodal fusion features are input into the context state modeler, and the context state modeler analyzes the context association information in the multimodal fusion features to generate context state information.
[0014] The context state information is input into the speech generation logic component, and the context state information is processed by the state and behavior controller, the speech generation strategy engine, the smooth transition controller and the speech text generator in the speech generation logic component to obtain the target speech text;
[0015] The voice command generator generates a target voice command based on the target voice text.
[0016] In an optional embodiment of the first aspect of the present invention, the step of processing the contextual state information through the state and behavior controller, the speech generation strategy engine, the smooth transition controller, and the speech-text generator in the speech generation logic component to obtain the target speech-text includes:
[0017] The state and behavior controller determines the target state and corresponding behavior instructions for speech generation based on the context state information.
[0018] The speech generation strategy engine selects a matching speech generation strategy from a preset strategy library based on the target state and the behavior instruction.
[0019] The smooth transition controller performs smooth processing on the state switching or parameter adjustment during the speech generation process according to the speech generation strategy, and outputs speech text generation instructions.
[0020] The speech text generator is controlled to generate the target speech text according to the speech text generation instructions.
[0021] In an optional embodiment of the first aspect of the present invention, the step of inputting the target speech instruction to the speech generation layer, and processing the target speech instruction through the speech parameter controller, speech synthesis scheduling component, and speech stream controller of the speech generation layer to obtain the target synthesized speech includes:
[0022] The target voice command is input into the voice parameter controller, and the voice generation parameters of the target voice command are adjusted by the voice parameter controller. The voice generation parameters include at least one of tone, speech rate and volume.
[0023] The parameter-tuned target speech command is input into the speech synthesis scheduling component, and the semantic segment splitter, the intonation controller, the TTS model engine and the post-processor in the speech synthesis scheduling component process the parameter-tuned target speech command to obtain intermediate synthesized speech;
[0024] The intermediate synthesized speech is input into the speech stream controller, and the playback stream of the intermediate synthesized speech is controlled by the speech stream controller to obtain the target synthesized speech.
[0025] In an optional embodiment of the first aspect of the present invention, the step of inputting the parameter-tuned target speech instruction into the speech synthesis scheduling component, and processing the parameter-tuned target speech instruction through the semantic segment splitter, the intonation controller, the TTS model engine, and the post-processor in the speech synthesis scheduling component to obtain intermediate synthesized speech includes:
[0026] The target speech command, after parameter tuning, is input into the semantic segment splitter to process the text portion, splitting the text portion of the target speech command into multiple semantic segments;
[0027] The target speech command after segmentation is input into the intonation controller, and the intonation of each semantic segment is controlled according to the speech generation parameters;
[0028] The initial speech is generated based on the target speech command after intonation adjustment using a TTS model engine;
[0029] The initial speech is filtered, denoised, gained, and optimized in terms of timbre by a post-processor to obtain intermediate synthesized speech.
[0030] In an optional embodiment of the first aspect of the present invention, the speech generation method further includes:
[0031] The processing of the multimodal data manager, the context and control layer, and the speech generation layer is monitored and scored in real time by a real-time monitoring and scoring system, and feedback information is generated.
[0032] The processing procedures of the multimodal data manager, the context and control layer, and the speech generation layer are dynamically adjusted based on the feedback information.
[0033] In an optional embodiment of the first aspect of the present invention, the dynamic adjustment of the processing procedures of the multimodal data manager, the context and control layer, and the speech generation layer based on the feedback information includes:
[0034] When the feedback information contains an interruption command, the current speech synthesis state information is saved, and when the interruption is resumed, the synthesis is resumed to the state before the interruption based on the state information.
[0035] A second aspect of the present invention provides a speech generation apparatus based on multimodal input, the speech generation apparatus based on multimodal input comprising:
[0036] The data receiving module is used to receive multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data.
[0037] The data fusion module is used to input the multimodal input data and the auxiliary input data into a preset multimodal data manager, perform fusion processing on the multimodal input data through the multimodal data manager, and combine the correlation features in the auxiliary input data to obtain multimodal fusion features;
[0038] The instruction generation module is used to input the multimodal fusion features into the context and control layer, and process the multimodal fusion features through the context state modeler, speech generation logic component and speech instruction generator in the context and control layer to obtain the target speech instruction. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller and a speech text generator.
[0039] The speech generation module is used to input the target speech command into the speech generation layer, and process the target speech command through the speech parameter controller, speech synthesis scheduling component and speech flow controller of the speech generation layer to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine and a post-processor.
[0040] The speech output module is used to output the target synthesized speech through the output layer, which includes at least one of local output, online download and playback, transcoding and embedding into other multimedia files, pushing to a third-party system, and asynchronous processing.
[0041] A third aspect of the present invention provides a speech generation device based on multimodal input, the speech generation device based on multimodal input comprising: a memory and at least one processor, the memory storing instructions, and the memory and the at least one processor being interconnected via a circuit;
[0042] The at least one processor invokes the instructions in the memory to cause the speech generation device based on multimodal input to perform the speech generation method based on multimodal input as described in any one of the first aspects of the invention.
[0043] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech generation method based on multimodal input as described in any one of the first aspects of the present invention.
[0044] Beneficial Effects: This invention provides a speech generation method and related equipment based on multimodal input. The method includes receiving multimodal input data and auxiliary input data, and inputting them into a multimodal data manager for fusion processing and extraction of associated features to obtain multimodal fusion features; inputting the multimodal fusion features into a context state modeler, a state and behavior controller, a speech generation strategy engine, a smooth transition controller, a speech-text generator, and a speech command generator to obtain a target speech command; inputting the target speech command into a speech parameter controller, a semantic segment splitter, a intonation controller, a TTS model engine, a post-processor, and a speech stream controller to obtain target synthesized speech; and outputting the target synthesized speech through an output layer in multiple ways. This invention integrates multimodal data and auxiliary data, and can generate speech based on contextual guidance of fusion features. The speech generation strategy is flexible and diverse, which can better meet user needs. Attached Figure Description
[0045] Figure 1 This is a schematic diagram illustrating one embodiment of the main steps of a speech generation method based on multimodal input according to the present invention;
[0046] Figure 2 This is a schematic diagram of an embodiment of the platform architecture of a speech generation method based on multimodal input according to the present invention;
[0047] Figure 3 This is a schematic diagram of an embodiment of a speech generation device based on multimodal input according to the present invention;
[0048] Figure 4 This is a schematic diagram of an embodiment of a speech generation device based on multimodal input according to the present invention. Detailed Implementation
[0049] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0050] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 The first aspect of this invention provides a speech generation method based on multimodal input, comprising:
[0051] S100. Receive multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data. The user operation data includes various operations performed by the user on the system or software, such as clicking on certain functions of the system or software. The real-time information capture data includes real-time data acquired through data acquisition devices (e.g., sensors and cameras). The extended input data includes extended data of various types of data. The preset parameters include preset limiting parameters for various types of data. The purpose of this invention is to achieve richer speech generation effects by fusing multimodal data.
[0052] S200. The multimodal input data and the auxiliary input data are input into a preset multimodal data manager. The multimodal data manager performs fusion processing on the multimodal input data and combines the correlation features in the auxiliary input data to obtain multimodal fusion features. In this invention, the process of receiving multimodal input data and auxiliary input data, as well as the multimodal data manager, constitute the input layer of the speech generation platform of this invention. In this invention, the multimodal data manager performs alignment, cleaning, and feature fusion processing on the multimodal input data (fusion methods include feature-level fusion, decision-level fusion, and hybrid-level fusion. Feature-level fusion directly concatenates feature vectors, decision-level fusion concatenates feature vectors after weighting, and hybrid-level fusion combines both fusion methods). Then, the multimodal fusion features are obtained by combining preset parameters and the correlation features of historical input data. The input layer of this invention can gather various heterogeneous signals such as text and multimedia information, user interaction operations, historical context, and preset parameters into the multimodal data manager to complete the initial cleaning, fusion, and formatting of the data, providing a unified, orderly, and traceable context benchmark for subsequent processing layers.
[0053] For example, the main data fusion process in step S200 may include: cross-modal semantic alignment, semantic role labeling of text input, extraction of core predicates and their associated entities; target detection and scene recognition of image input, extraction of visual entities and scene categories; calculation of semantic similarity between text entities and visual entities through a cross-modal attention mechanism, generating a text-image alignment weight matrix.
[0054] Ambient sound and user operation are analyzed collaboratively. Acoustic features are extracted from ambient sound input to identify environmental event types; behavioral patterns are analyzed from user operation input to identify operation intentions; and the timestamps of environmental events and user operations are aligned based on the Dynamic Time Warping (DTW) algorithm to construct an ambient sound-user operation association graph.
[0055] Historical data is dynamically adapted by extracting historical scene types, user preferences, and incomplete instructions from the historical input of auxiliary input data; the current multimodal input and historical data are time-series modeled by a gated recurrent unit (GRU) to output historical adaptation weights.
[0056] Multimodal feature fusion is performed based on the text-image alignment weight matrix, ambient sound-user operation association graph, and historical adaptation weights. A feature concatenation and gating fusion mechanism is used to generate multimodal fused features, wherein: text and image features are concatenated according to alignment weights; ambient sound and user operation features are aggregated through graph convolution through the association graph; and historical adaptation weights control the fusion ratio between current features and historical features.
[0057] The associated features are structured and output, encoding the multimodal fusion features into three layers of structured features: spatiotemporal feature layer: timestamps and spatial coordinates of labeled features; semantic cluster layer: features with clustered semantic associations; and behavior chain layer: user operation behaviors are linked together in time sequence.
[0058] S300. The multimodal fusion features are input into the context and control layer. The multimodal fusion features are processed by the context state modeler, speech generation logic component, and speech command generator in the context and control layer to obtain the target speech command. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller, and a speech text generator. In this invention, the context and control layer is the perception and scheduling core of the intelligent decision-making of the speech synthesis platform of this invention. The context state modeler extracts user intent, environmental features, and interaction history from the aggregated multi-source inputs through a language big model and rule engine to construct a real-time updated "dialogue state profile".
[0059] The state and behavior controller, based on the current state profile and integrating a strategy engine and sentiment analysis, determines the style, emotional tone, and rhythm of the voice output, deciding "what to say" and "in what tone." The speech-to-text generator transforms the semantic intent output by the controller into a fluent and natural text, balancing grammatical correctness and elegant writing, laying the textual foundation for downstream TTS. The smooth transition and instruction generation module automatically generates synthesized instructions that can be recognized by the TTS engine through optimization of the connection between preceding and following sentences, ensuring overall auditory coherence and natural emotional transitions.
[0060] In an optional embodiment of the first aspect of the present invention, step S300 includes:
[0061] S301. The multimodal fusion feature is input into the context state modeler, and the context state modeler analyzes the context association information in the multimodal fusion feature to generate context state information. In this invention, the context association information includes the time series association between the multimodal fusion feature and historical input data, the semantic matching association between text input semantics and image visual content, and the context association between user operation behavior and environmental sound features. The context state modeler can use the Transformer model to perform time series modeling on the context association information to generate context state information containing the current scene type, user intent category, and sentiment tendency.
[0062] S302. The context state information is input into the speech generation logic component. The context state information is processed by the state and behavior controller, the speech generation strategy engine, the smooth transition controller, and the speech text generator in the speech generation logic component to obtain the target speech text. Specifically, an exemplary implementation of this step may include: determining the target state and corresponding behavior instruction for speech generation based on the context state information by the state and behavior controller; selecting a matching speech generation strategy from a preset strategy library based on the target state and the behavior instruction by the speech generation strategy engine; smoothing the state switching or parameter adjustment during the speech generation process by the smooth transition controller based on the speech generation strategy, and outputting a speech text generation instruction; and controlling the speech text generator to generate the target speech text according to the speech text generation instruction.
[0063] S303. The voice command generator generates a target voice command based on the target voice text. In this invention, after receiving the target voice text, the voice command generator performs semantic dependency analysis on the target voice text, extracts the three elements of the command: the command body, the command action, and the command parameters, and generates a target voice command that conforms to the interface specification.
[0064] S400. The target speech command is input to the speech generation layer. The speech parameter controller, speech synthesis scheduling component, and speech stream controller of the speech generation layer process the target speech command to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine, and a post-processor. In this invention, the speech generation layer mainly realizes the synthesis of sound data and post-processing fine-tuning. It is the production site for the speech file to be heard. The speech parameter controller can finely adjust parameters such as speech rate, pitch, and stress to achieve diverse sound images. The semantic segment splitter can intelligently segment long texts to ensure that each syllable falls in the most natural position. The TTS model engine can deeply call advanced synthesis models such as FastSpeech and VITS to convert text into high-fidelity audio. The post-processor can apply techniques such as filtering, noise reduction, and gain control to polish the generated audio and eliminate the mechanical feel. The speech stream controller can reasonably schedule the playback order and resource usage in concurrent and multi-output scenarios to achieve second-level response and seamless connection.
[0065] In an optional embodiment of the first aspect of the present invention, step S400 includes:
[0066] 401. Input the target voice command into the voice parameter controller, and adjust the voice generation parameters of the target voice command through the voice parameter controller. The voice generation parameters include at least one of tone, speech rate, and volume. The process of adjusting the voice generation parameters of the target voice command by the voice parameter controller includes: extracting the tone reference range, speech rate reference value, and volume reference range of the corresponding scene from a preset scene-parameter mapping table according to the scene type in the command parameters; and making targeted fine adjustments to the tone reference range, speech rate reference value, and volume reference range according to the user intent category in the command parameters.
[0067] 402. The parameter-tuned target speech command is input into the speech synthesis scheduling component. The semantic segment splitter, intonation controller, TTS model engine, and post-processor in the speech synthesis scheduling component process the parameter-tuned target speech command to obtain intermediate synthesized speech. An exemplary implementation of this step may include: inputting the parameter-tuned target speech command into the semantic segment splitter to process the text portion, splitting the text portion of the target speech command into multiple semantic segments; inputting the segmented target speech command into the intonation controller, controlling the intonation of each semantic segment according to the speech generation parameters; generating initial speech based on the intonation-tuned target speech command through the TTS model engine; and filtering, denoising, gaining, and timbre optimization of the initial speech through the post-processor to obtain intermediate synthesized speech.
[0068] 403. The intermediate synthesized speech is input into the speech stream controller, and the playback stream of the intermediate synthesized speech is controlled by the speech stream controller to obtain the target synthesized speech. The playback stream control methods include interruption recovery, rollback switching, and real-time stream adjustment.
[0069] S500. The target synthesized speech is output through the output layer, which includes at least one of the following: local output, online download and playback, transcoding and embedding into other multimedia files, push to a third-party system, and asynchronous processing. The speech output layer of this invention enables flexible speech delivery in multiple scenarios. At the output layer, the platform can provide multiple delivery paths in a highly decoupled manner: Local playback: Supports real-time playback on devices such as mobile phones, smart speakers, and headphones. Online download and embedding: Can generate downloadable files in multiple formats, or seamlessly embed audio into multimedia content such as videos, PPTs, and games. Third-party push: Pushes the generated results to third-party business systems via API or message queues. Asynchronous task processing: For scenarios that do not require real-time response, the system can queue and execute tasks in the background, scheduling appropriate resources to achieve optimal cost-efficiency.
[0070] In an optional embodiment of the first aspect of the present invention, the speech generation method further includes: real-time monitoring and quality scoring of the processing of the multimodal data manager, the context and control layer, and the speech generation layer through a real-time monitoring and scoring system, generating feedback information; and dynamically adjusting the processing of the multimodal data manager, the context and control layer, and the speech generation layer based on the feedback information. In an exemplary scenario of this step, when the feedback information contains an interruption command, the current speech synthesis state information is saved, and upon resumption of the interruption, the synthesis is reverted to the state before the interruption based on the state information. Throughout the entire context awareness and control process of the present invention, a multi-dimensional monitoring and scoring system can be run to ensure the continuity of interaction and output quality. Its core responsibilities include user-driven immediate correction or interruption and real-time scoring and reconstruction of content quality.
[0071] In summary, the platform architecture of the speech generation scheme based on multimodal input in this invention can be as follows: Figure 2 As shown, the main features of the speech generation scheme based on multimodal input in this invention are: (1) Multimodal input aggregation: The system simultaneously accesses various heterogeneous information such as text, speech segments, images, ambient sounds, sensor data, user operations, and historical context, and completes unified cleaning, formatting, and management in the "multimodal data manager". (2) Real-time monitoring correction mechanism: It has two fault-tolerant paths built-in: user-driven interruption and real-time content quality scoring, allowing users to interrupt and readjust the process at any time; in the summary output scenario, the generated results are automatically scored and low scores are rolled back to ensure overall quality. (3) High-fidelity speech synthesis and post-processing: The speech rate, tone, and emotion are finely adjusted by the "speech parameter controller", and then the initial synthesis is completed by advanced TTS models (such as FastSpeech and VITS). Finally, after post-processing such as "semantic segment splitter", "filtering, noise reduction, and gain control", a natural, fluent, and high-quality audio stream is output. (4) Flexible delivery in multiple scenarios: Audio results can be played locally in real time, or generated as downloadable files, embedded in multimedia, pushed to third-party systems, or processed in batches as asynchronous background tasks, achieving end-to-end full-scenario coverage.
[0072] Based on the above-mentioned main features of the speech generation scheme of the present invention, the speech generation scheme of the present invention can bring at least the following beneficial effects:
[0073] 1. Highly personalized auditory experience: Through the dynamic fusion of multimodal (expandable) information, the system can adjust the tone, rhythm and emotional color of the voice in real time, making each broadcast seem "customized" to fit the user's current mood and scene needs.
[0074] 2. Continuity of interactive response: Users can initiate ratings or the system can automatically trigger correction commands. The system can capture and reset the generation process in real time to avoid "stuck" or "awkward" moments. At the same time, the smooth transition module ensures that the new and old content are connected naturally without any abruptness.
[0075] 3. Intelligent control of output quality: Through real-time scoring and rollback mechanisms, a quality lock is added to "summary" or "batch" generation scenarios. When the generated content does not meet the preset standards, it can be automatically reconstructed, greatly reducing the risk of information omission and semantic deviation.
[0076] See Figure 3 A second aspect of the present invention provides a speech generation apparatus based on multimodal input, the speech generation apparatus based on multimodal input comprising:
[0077] Data receiving module 10 is used to receive multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data.
[0078] The data fusion module 20 is used to input the multimodal input data and the auxiliary input data into a preset multimodal data manager, perform fusion processing on the multimodal input data through the multimodal data manager, and combine the correlation features in the auxiliary input data to obtain multimodal fusion features;
[0079] The instruction generation module 30 is used to input the multimodal fusion features into the context and control layer, and process the multimodal fusion features through the context state modeler, speech generation logic component and speech instruction generator in the context and control layer to obtain the target speech instruction. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller and a speech text generator.
[0080] The speech generation module 40 is used to input the target speech command to the speech generation layer, and process the target speech command through the speech parameter controller, speech synthesis scheduling component and speech flow controller of the speech generation layer to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine and a post-processor.
[0081] The speech output module 50 is used to output the target synthesized speech through the output layer, which includes at least one of local output, online download and playback, transcoding and embedding into other multimedia files, pushing to a third-party system, and asynchronous processing.
[0082] In an optional embodiment of the second aspect of the present invention, the instruction generation module includes:
[0083] The state modeling processing unit is used to input the multimodal fusion features into the context state modeler, and analyze the context association information in the multimodal fusion features through the context state modeler to generate context state information;
[0084] The speech-to-text generation unit is used to input the context state information into the speech generation logic component, and process the context state information through the state and behavior controller, the speech generation strategy engine, the smooth transition controller and the speech-to-text generator in the speech generation logic component to obtain the target speech-to-text.
[0085] A voice command generation unit is used to generate target voice commands based on the target speech text through the voice command generator.
[0086] In an optional embodiment of the second aspect of the present invention, the speech-text generation unit includes:
[0087] The state / instruction determination subunit is used to determine the target state and corresponding behavior instruction for speech generation based on the context state information through the state and behavior controller.
[0088] A strategy matching subunit is generated, which is used to select a matching speech generation strategy from a preset strategy library based on the target state and the behavior instruction through the speech generation strategy engine;
[0089] The strategy smoothing processing subunit is used to smooth the state switching or parameter adjustment in the speech generation process according to the speech generation strategy through the smooth transition controller, and output the speech text generation instruction.
[0090] The instruction text conversion subunit is used to control the speech text generator to generate target speech text according to the speech text generation instruction.
[0091] In an optional embodiment of the second aspect of the present invention, the speech generation module includes:
[0092] A voice parameter acquisition unit is used to input the target voice command into the voice parameter controller and adjust the voice generation parameters of the target voice command through the voice parameter controller. The voice generation parameters include at least one of tone, speech rate and volume.
[0093] An intermediate speech synthesis unit is used to input the parameter-tuned target speech instruction into the speech synthesis scheduling component, and process the parameter-tuned target speech instruction through the semantic segment splitter, the intonation controller, the TTS model engine and the post-processor in the speech synthesis scheduling component to obtain intermediate synthesized speech;
[0094] The target speech synthesis unit is used to input the intermediate synthesized speech into the speech stream controller, and control the playback stream of the intermediate synthesized speech through the speech stream controller to obtain the target synthesized speech.
[0095] In an optional embodiment of a first aspect of the present invention, the intermediate speech synthesis unit includes:
[0096] The paragraph splitting subunit is used to input the parameter-tuned target speech command into the semantic paragraph splitter to process the text part and split the text part of the target speech command into multiple semantic paragraphs;
[0097] The intonation control subunit is used to input the target speech instruction after paragraph segmentation into the intonation controller, and control the intonation of each semantic paragraph according to the speech generation parameters;
[0098] The instruction speech conversion subunit is used to generate initial speech based on the target speech instruction after intonation adjustment through the TTS model engine;
[0099] The speech post-processing subunit is used to filter, reduce noise, increase gain, and optimize timbre of the initial speech through a post-processor to obtain intermediate synthesized speech.
[0100] In an optional embodiment of the second aspect of the present invention, the speech generation apparatus further includes:
[0101] The process feedback module is used to monitor and score the processing of the multimodal data manager, the context and control layer, and the speech generation layer in real time through a real-time monitoring and scoring system, and generate feedback information.
[0102] The process adjustment module is used to dynamically adjust the processing of the multimodal data manager, the context and control layer, and the speech generation layer based on the feedback information.
[0103] In an optional embodiment of the second aspect of the present invention, the process adjustment module includes:
[0104] An interruption instruction processing unit is used to save the current speech synthesis state information when the feedback information contains an interruption instruction, and to revert to the state before the interruption and continue synthesis when the interruption is resumed.
[0105] Figure 4 This is a schematic diagram of a speech generation device based on multimodal input provided by an embodiment of the present invention. This speech generation device can vary considerably depending on its configuration or performance, and may include one or more processors 60 (central processing units, CPUs) (e.g., one or more processors) and a memory 70, and one or more storage media 80 (e.g., one or more mass storage devices) for storing applications or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the speech generation device based on multimodal input. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the speech generation device based on multimodal input.
[0106] The speech generation device based on multimodal input of this invention may further include one or more power supplies 90, one or more wired or wireless network interfaces 100, one or more input / output interfaces 110, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4 The illustrated structure of a speech generation device based on multimodal input does not constitute a limitation on speech generation devices based on multimodal input. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0107] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the speech generation method based on multimodal input.
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system or system / unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech generation method based on multimodal input, characterized in that, include: The system receives multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data. The multimodal input data and the auxiliary input data are input into a preset multimodal data manager. The multimodal input data is fused by the multimodal data manager, and the multimodal fusion features are obtained by combining the correlation features in the auxiliary input data. The multimodal fusion features are input into the context and control layer. The multimodal fusion features are processed by the context state modeler, speech generation logic component and speech command generator in the context and control layer to obtain the target speech command. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller and a speech text generator. The target speech command is input to the speech generation layer, and the speech parameter controller, speech synthesis scheduling component and speech stream controller of the speech generation layer process the target speech command to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine and a post-processor. The target synthesized speech is output through an output layer, which includes at least one of the following: local output, online download and playback, transcoding and embedding into other multimedia files, pushing to a third-party system, and asynchronous processing. The process of inputting the multimodal fusion features into the context and control layer, and processing the multimodal fusion features through the context state modeler, speech generation logic component, and speech command generator in the context and control layer to obtain the target speech command includes: The multimodal fusion features are input into the context state modeler, and the context state modeler analyzes the context association information in the multimodal fusion features to generate context state information. The context state information is input into the speech generation logic component, and the context state information is processed by the state and behavior controller, the speech generation strategy engine, the smooth transition controller and the speech text generator in the speech generation logic component to obtain the target speech text; The voice command generator generates a target voice command based on the target voice text. The process of processing the contextual state information through the state and behavior controller, the speech generation strategy engine, and the speech-text generator in the speech generation logic component to obtain the target speech-text includes: The state and behavior controller determines the target state and corresponding behavior instructions for speech generation based on the context state information. The speech generation strategy engine selects a matching speech generation strategy from a preset strategy library based on the target state and the behavior instruction. The smooth transition controller performs smooth processing on the state switching or parameter adjustment during the speech generation process according to the speech generation strategy, and outputs speech text generation instructions. The speech text generator is controlled to generate the target speech text according to the speech text generation instructions.
2. The speech generation method based on multimodal input according to claim 1, characterized in that, The step of inputting the target speech command to the speech generation layer, and processing the target speech command through the speech parameter controller, speech synthesis scheduling component, and speech stream controller of the speech generation layer to obtain the target synthesized speech includes: The target voice command is input into the voice parameter controller, and the voice generation parameters of the target voice command are adjusted by the voice parameter controller. The voice generation parameters include at least one of tone, speech rate and volume. The parameter-tuned target speech command is input into the speech synthesis scheduling component, and the semantic segment splitter, the intonation controller, the TTS model engine and the post-processor in the speech synthesis scheduling component process the parameter-tuned target speech command to obtain intermediate synthesized speech; The intermediate synthesized speech is input into the speech stream controller, and the playback stream of the intermediate synthesized speech is controlled by the speech stream controller to obtain the target synthesized speech.
3. The speech generation method based on multimodal input according to claim 2, characterized in that, The step of inputting the parameter-tuned target speech command into the speech synthesis scheduling component, and processing the parameter-tuned target speech command through the semantic segment splitter, the intonation controller, the TTS model engine, and the post-processor in the speech synthesis scheduling component to obtain intermediate synthesized speech includes: The target speech command, after parameter tuning, is input into the semantic segment splitter to process the text portion, splitting the text portion of the target speech command into multiple semantic segments; The target speech command after segmentation is input into the intonation controller, and the intonation of each semantic segment is controlled according to the speech generation parameters; The initial speech is generated based on the target speech command after intonation adjustment using a TTS model engine; The initial speech is filtered, denoised, gained, and optimized in terms of timbre by a post-processor to obtain intermediate synthesized speech.
4. The speech generation method based on multimodal input according to claim 1, characterized in that, The speech generation method further includes: The processing of the multimodal data manager, the context and control layer, and the speech generation layer is monitored and scored in real time by a real-time monitoring and scoring system, and feedback information is generated. The processing procedures of the multimodal data manager, the context and control layer, and the speech generation layer are dynamically adjusted based on the feedback information.
5. The speech generation method based on multimodal input according to claim 4, characterized in that, The dynamic adjustment of the processing procedures of the multimodal data manager, the context and control layer, and the speech generation layer based on the feedback information includes: When the feedback information contains an interruption command, the current speech synthesis state information is saved, and when the interruption is resumed, the synthesis is resumed to the state before the interruption based on the state information.
6. A speech generation device based on multimodal input, characterized in that, The speech generation device based on multimodal input includes: The data receiving module is used to receive multimodal input data and auxiliary input data. The multimodal input data includes at least two of the following: text, image, ambient sound, user operation data, real-time information capture data, and extended input data. The auxiliary input data includes preset parameters and historical input data. The data fusion module is used to input the multimodal input data and the auxiliary input data into a preset multimodal data manager, perform fusion processing on the multimodal input data through the multimodal data manager, and combine the correlation features in the auxiliary input data to obtain multimodal fusion features; The instruction generation module is used to input the multimodal fusion features into the context and control layer, and process the multimodal fusion features through the context state modeler, speech generation logic component and speech instruction generator in the context and control layer to obtain the target speech instruction. The speech generation logic component includes a state and behavior controller, a speech generation strategy engine, a smooth transition controller and a speech text generator. The speech generation module is used to input the target speech command into the speech generation layer, and process the target speech command through the speech parameter controller, speech synthesis scheduling component and speech flow controller of the speech generation layer to obtain the target synthesized speech. The speech synthesis scheduling component includes a semantic segment splitter, a intonation controller, a TTS model engine and a post-processor. The speech output module is used to output the target synthesized speech through the output layer, and the output layer includes at least one of local output, online download and playback, transcoding and embedding into other multimedia files, pushing to a third-party system, and asynchronous processing; The instruction generation module includes: The state modeling processing unit is used to input the multimodal fusion features into the context state modeler, and analyze the context association information in the multimodal fusion features through the context state modeler to generate context state information; The speech-to-text generation unit is used to input the context state information into the speech generation logic component, and process the context state information through the state and behavior controller, the speech generation strategy engine, the smooth transition controller and the speech-to-text generator in the speech generation logic component to obtain the target speech-to-text. A voice instruction generation unit is used to generate a target voice instruction based on the target voice text through the voice instruction generator; The speech-to-text generation unit includes: The state / instruction determination subunit is used to determine the target state and corresponding behavior instruction for speech generation based on the context state information through the state and behavior controller. A strategy matching subunit is generated, which is used to select a matching speech generation strategy from a preset strategy library based on the target state and the behavior instruction through the speech generation strategy engine; The strategy smoothing processing subunit is used to smooth the state switching or parameter adjustment in the speech generation process according to the speech generation strategy through the smooth transition controller, and output the speech text generation instruction. The instruction text conversion subunit is used to control the speech text generator to generate target speech text according to the speech text generation instruction.
7. A speech generation device based on multimodal input, characterized in that, The speech generation device based on multimodal input includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor invokes the instructions in the memory to cause the speech generation device based on multimodal input to perform the speech generation method based on multimodal input as described in any one of claims 1-5.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the speech generation method based on multimodal input as described in any one of claims 1-5.
Citation Information
Patent Citations
Data conversion method and device and electronic equipment
CN118918877A
Method and device for converting text into voice, equipment and storage medium
CN119517004A
Multi-modal voice interaction method based on large model, electronic equipment and storage medium
CN119559946A