Voice processing method and device based on artificial intelligence, computer equipment and medium

Through an AI-based multimodal fusion architecture and the use of technologies such as voice encoders, adapters, and decoders, the problems of high latency and low quality in traditional voice interaction systems have been solved, achieving efficient and natural voice interaction and improving user experience and business efficiency.

CN120690194APending Publication Date: 2025-09-23CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510845854.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional voice interaction systems have problems with high interaction delay and low voice response quality, resulting in a poor user experience.

Method used

It adopts an AI-based multimodal fusion architecture, including pre-trained speech encoders, speech adapters, large language models, and streaming speech decoders, to achieve seamless conversion from speech to text and then back to speech, and generates high-quality speech feedback through feature extraction, adjustment, inference, decoding, and optimization processing.

Benefits of technology

It achieves real-time and natural voice interaction, improves processing efficiency and the quality of generated voice, and enhances customer service experience and business processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690194A_ABST
    Figure CN120690194A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an artificial intelligence-based voice processing method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a voice instruction inputted by a user through voice equipment; performing feature extraction on the voice instruction based on a voice encoder to obtain voice features; performing adjustment processing on the voice feature based on a voice adapter to obtain a target voice feature; reasoning the target speech features based on a large language model to generate a target text; decoding the target text based on a voice decoder to obtain a reply voice; performing optimization processing on the reply voice based on a quality optimization strategy to obtain a target reply voice; and transmitting the target reply voice to the voice equipment based on the playing control strategy. In addition, the invention also relates to a block chain technology, and the target reply voice can be stored in the block chain. The voice interaction processing method and device can be applied to voice interaction scenes in the financial field and the medical field, and the voice interaction processing efficiency is effectively improved through the voice interaction processing method and device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as financial technology and digital medicine, and in particular to artificial intelligence-based speech processing methods, devices, computer equipment and storage media. Background Art

[0002] In traditional customer service industries, particularly in finance and healthcare, voice interaction over the phone is a common service method. In finance, for example, when an insurance customer inquires about their insurance product, traditional voice interaction systems typically first convert the customer's voice instructions into text using automatic speech recognition (ASR) technology, then perform business logic processing based on the text, and finally use text-to-speech (TTS) technology to convert the processed results into voice feedback for the customer. Similarly, in healthcare, when a patient inquires about registration procedures or precautions regarding a disease over the phone with a hospital, traditional ASR and TTS technologies are also relied upon for voice interaction.

[0003] However, traditional ASR and TTS technologies have significant flaws. For one thing, traditional technologies often employ a serial processing approach, with speech recognition, business processing, and speech synthesis performed sequentially, resulting in high overall interaction latency. In financial consulting scenarios, after a customer asks a question, they may have to wait several seconds or even longer for a response, severely impacting service efficiency. Furthermore, the speech generated by traditional TTS technology sounds mechanical and unnatural, making it difficult to simulate the intonation, pauses, and emotional expression of real human speech, resulting in poor quality of generated voice responses. In medical consultation scenarios, patients can easily experience communication barriers when hearing stiff, unnatural voice responses, reducing their trust in the service.

[0004] These shortcomings have resulted in poor user experiences for traditional voice interaction systems in related fields, failing to meet customer demands for efficient and natural interactions. Therefore, a new voice interaction technology is urgently needed that can reduce interaction latency to improve user experience and, in turn, enhance the quality of customer service in related industries. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to propose a voice processing method, device, computer equipment and storage medium based on artificial intelligence to solve the technical problems that the existing voice interaction mode has low processing efficiency due to high interaction delay and poor quality of generated voice response.

[0006] In a first aspect, a speech processing method based on artificial intelligence is provided, comprising:

[0007] Receive voice commands input by the user through a preset voice device;

[0008] Extracting features of the voice command based on a preset voice encoder to obtain corresponding voice features;

[0009] Adjusting and processing the voice features based on a preset voice adapter to obtain corresponding target voice features;

[0010] Performing inference processing on the target speech features based on a preset large language model to generate corresponding target text;

[0011] Decoding the target text based on a preset voice decoder to obtain a corresponding reply voice;

[0012] Optimizing the reply voice based on a preset quality optimization strategy to obtain a corresponding target reply voice;

[0013] Based on a preset playback control strategy, the target reply voice is transmitted to the voice device.

[0014] In a second aspect, a speech processing device based on artificial intelligence is provided, comprising:

[0015] A receiving module, configured to receive voice commands input by a user through a preset voice device;

[0016] An extraction module, configured to extract features from the voice command based on a preset voice encoder to obtain corresponding voice features;

[0017] An adjustment module, configured to adjust the voice features based on a preset voice adapter to obtain corresponding target voice features;

[0018] An inference module, configured to perform inference processing on the target speech features based on a preset large language model to generate a corresponding target text;

[0019] A decoding module, configured to decode the target text based on a preset voice decoder to obtain a corresponding reply voice;

[0020] An optimization module, configured to optimize the reply speech based on a preset quality optimization strategy to obtain a corresponding target reply speech;

[0021] The transmission module is used to transmit the target reply voice to the voice device based on a preset playback control strategy.

[0022] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned artificial intelligence-based speech processing method when executing the computer program.

[0023] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned artificial intelligence-based speech processing method are implemented.

[0024] In the solution implemented by the above-mentioned artificial intelligence-based speech processing method, device, computer equipment and storage medium, first, a voice command input by a user through a preset voice device is received; then, based on a preset voice encoder, feature extraction is performed on the voice command to obtain corresponding voice features; then, based on a preset voice adapter, the voice features are adjusted and processed to obtain corresponding target voice features; and based on a preset large language model, the target voice features are inferred and processed to generate corresponding target text; subsequently, the target text is decoded and processed based on a preset voice decoder to obtain a corresponding reply voice; further, the reply voice is optimized based on a preset quality optimization strategy to obtain a corresponding target reply voice; finally, based on a preset playback control strategy, the target reply voice is transmitted to the voice device. Based on the above processing flow, this application, by combining the use of a voice encoder, a voice adapter, a large language model and a voice decoder, can communicate with the user in real time and naturally, accurately understand the voice commands input by the user and provide high-quality voice feedback, effectively improve the processing efficiency of voice interaction, ensure the quality of the generated target reply voice, and thus help improve customer service experience and business processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0027] Figure 2 is a flow chart of an embodiment of an artificial intelligence-based speech processing method according to the present application;

[0028] Figure 3 is a structural diagram of an embodiment of an artificial intelligence-based speech processing device according to the present application;

[0029] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0031] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0032] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0033] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0034] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0035] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0036] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0037] It should be noted that the artificial intelligence-based speech processing method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the artificial intelligence-based speech processing device is generally set in the server / terminal device.

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0039] Continue to refer Figure 2 , showing a flow chart of an embodiment of the speech processing method based on artificial intelligence according to the present application. Depending on different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted. The speech processing method based on artificial intelligence provided in the embodiment of the present application can be applied to any scenario that requires voice interaction processing, and the speech processing method based on artificial intelligence can be applied to products in these scenarios, for example, voice interaction scenarios in the financial field (such as customer service, claims processing, and insurance consulting business scenarios) and the medical field (such as online consultation, medication consultation, and other business scenarios). The speech processing method based on artificial intelligence comprises the following steps:

[0040] Step S201: receiving a voice command input by a user through a preset voice device.

[0041] In this embodiment, the electronic device (e.g. Figure 1The server / terminal device shown in the figure) can obtain the voice instructions input by the user through the voice device through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wi deband) connection, and other wireless connection methods currently known or to be developed in the future. The execution subject of this application is specifically a voice processing system, which can be referred to as the system for short. The above-mentioned voice device may be a device such as a landline phone, a smart phone, a smart speaker, etc. The user can input the corresponding voice instructions through the above-mentioned voice device according to the actual interaction needs. Among them, the system is a multimodal fusion architecture that integrates a pre-trained voice encoder, a voice adapter, a large language model and a streaming voice decoder to achieve seamless conversion from speech to text and then to speech. This multimodal fusion architecture not only improves the interaction efficiency of the system, but also significantly reduces the response delay of the system by eliminating the need for intermediate text transcription.

[0042] In addition, this application can be applied to question-and-answer processing scenarios in the financial and medical fields. For example, in the financial field, a user calls an insurance company's customer service number and asks the system relevant questions or requests, such as "I want to know about the car insurance claims process." In the medical field, a user calls a hospital's customer service number and asks the system relevant questions or requests, such as "I want to consult daily dietary management recommendations for diabetic patients."

[0043] Step S202: extracting features from the voice command based on a preset voice encoder to obtain corresponding voice features.

[0044] In this embodiment, a pre-trained speech encoder (such as Wav2Vec2.0 or HuBERT) can be used to extract features from the received raw speech signal (voice command). This speech encoder uses a deep learning model to convert the speech waveform into a high-dimensional vector representation (i.e., speech features). These vectors capture the key characteristics of the speech signal (such as phonemes, intonation, rhythm, etc.) while reducing the data dimensionality for subsequent processing.

[0045] Step S203: adjusting the voice features based on a preset voice adapter to obtain corresponding target voice features.

[0046] In this embodiment, the specific implementation process of adjusting the voice features based on the preset voice adapter to obtain the corresponding target voice features will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0047] Step S204: performing inference processing on the target speech features based on a preset large language model to generate a corresponding target text.

[0048] In this embodiment, the specific implementation process of inferring the target speech features based on the preset large-scale language model to generate the corresponding target text will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0049] Step S205: decoding the target text based on a preset voice decoder to obtain a corresponding reply voice.

[0050] In this embodiment, the above-mentioned specific implementation process of decoding the target text based on the preset voice decoder to obtain the corresponding reply voice will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0051] Step S206: Optimize the reply voice based on a preset quality optimization strategy to obtain a corresponding target reply voice.

[0052] In this embodiment, the above-mentioned specific implementation process of optimizing the reply voice based on the preset quality optimization strategy to obtain the corresponding target reply voice will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0053] Step S207: Based on a preset playback control strategy, the target reply voice is transmitted to the voice device.

[0054] In this embodiment, the specific implementation process of transmitting the target reply voice to the voice device based on the preset playback control strategy will be further described in detail in subsequent specific embodiments of this application and will not be elaborated on here.

[0055] This application first receives the voice command input by the user through a preset voice device; then extracts the features of the voice command based on the preset voice encoder to obtain the corresponding voice features; then adjusts and processes the voice features based on the preset voice adapter to obtain the corresponding target voice features; and infers and processes the target voice features based on the preset large language model to generate the corresponding target text; subsequently, decodes and processes the target text based on the preset voice decoder to obtain the corresponding reply voice; further optimizes and processes the reply voice based on the preset quality optimization strategy to obtain the corresponding target reply voice; finally, based on the preset playback control strategy, transmits the target reply voice to the voice device. Based on the above processing flow, this application can communicate with users in real time and naturally by combining the use of voice encoders, voice adapters, large language models, and voice decoders, accurately understand the voice commands input by users, and provide high-quality voice feedback, effectively improving the processing efficiency of voice interaction, ensuring the quality of the generated target reply voice, and thus helping to improve customer service experience and business processing efficiency.

[0056] In some optional implementations, step S204 includes the following steps:

[0057] The target speech features are inferred and processed based on the large language model to obtain corresponding multiple candidate text data.

[0058] In this embodiment, the selection of the aforementioned large-scale language model is not specifically limited and can be determined based on actual business needs. For example, the LLaMA-Omni model can be used. The aforementioned inference processing based on the large-scale language model includes the following steps: the input layer of the large-scale language model receives speech features that have undergone speech encoding and adaptive conversion. These features are represented as tensors and contain key information about the speech signal. The large-scale language model then performs in-depth analysis of the input speech features using a self-attention mechanism. This self-attention mechanism allows the model to focus on all other positions in the input sequence while processing features at each position, thereby capturing long-range dependencies and contextual information in the speech content. After processing by the self-attention mechanism, the features undergo a nonlinear transformation using a feedforward neural network (FFN). The FFN consists of multiple fully connected layers that further extract and integrate high-level semantic information from the speech features. Subsequently, the large-scale language model combines its pre-trained knowledge with the current input speech features to deeply understand and infer the user's intent and needs. For example, it can recognize that the user is inquiring about the car insurance claim process, rather than other insurance products. Finally, the decoder of the large-scale language model generates a corresponding text response word by word based on the understood contextual information. The decoder uses an autoregressive approach, generating a token at each step until a complete sentence is generated. The generated text response closely matches the user's voice command semantically. For example, when a user inquires about the car insurance claim process, the model generates a text response such as "The car insurance claim process is as follows: First, you need to report the accident to the insurance company within 48 hours of the accident..." This ensures that the information is accurate and easy to understand.

[0059] For example, when a user inquires about daily dietary management recommendations for diabetic patients, the model generates a text response such as "Daily dietary management recommendations for diabetic patients are as follows: Control carbohydrate intake: Give priority to low GI (glycemic index) foods, such as whole grains and oats, and avoid refined sugar and white rice. Balanced diet: Each meal is paired with vegetables, high-quality protein (such as fish and beans) and healthy fats (such as olive oil). Regular and quantitative meals: Eat small meals frequently to avoid drastic fluctuations in blood sugar. It is recommended to eat 5-6 meals a day. Monitor blood sugar: Adjust the diet plan according to changes in blood sugar after meals, and consult a doctor or nutritionist if necessary." to ensure that the information is accurate and easy to understand.

[0060] Call the preset speech model scoring module.

[0061] In this embodiment, when a large language model generates text, the large language model can simultaneously generate multiple candidate text segments (which may be referred to as candidate texts for short), and each candidate text segment can be evaluated by the speech model scoring module.

[0062] Generate scores for each of the candidate text data based on the speech model scoring module.

[0063] In this embodiment, the text scoring process of the speech model scoring module includes: scoring mechanism initialization, text segment scoring, and fluency threshold determination. Specifically, scoring mechanism initialization includes: the system loading a pre-trained language model scoring module (such as GPT-2, BERT, etc.), which is specifically used to evaluate the fluency and naturalness of text. The parameters of the speech model scoring module, such as the scoring threshold and maximum generation length, are configured according to the specific application scenario and requirements to ensure that the scoring mechanism meets the requirements of the actual application.

[0064] Text segment scoring involves the speech model scoring module performing real-time scoring on each newly generated token during text generation using a large language model. Scoring is based on the token's consistency with the context, grammatical correctness, and semantic plausibility. The speech model scoring module not only assesses the quality of individual tokens but also analyzes the contextual coherence of the entire text segment, ensuring that the generated text is logically coherent and seamless.

[0065] Fluency threshold assessment involves the system dynamically determining whether a generated text segment meets fluency requirements based on a preset fluency threshold. This threshold can be adjusted based on actual needs to balance generation speed and quality. If a text segment's score falls below the threshold, the system triggers a regeneration or adjustment mechanism, such as backtracking, replacing tokens, or adjusting the generation strategy to improve text fluency.

[0066] The designated candidate text data with the highest score is selected from all the candidate text data.

[0067] In this embodiment, the system selects the one with the highest score from multiple candidate segments (ie, candidate text data) as the final output, which can ensure that the final generated target text is optimal in terms of fluency and naturalness.

[0068] The designated candidate text data is used as the target text.

[0069] In this embodiment, the system also supports a user feedback mechanism, which can integrate user evaluations of generated text (such as satisfaction ratings) into the speech model scoring module to further optimize the effectiveness of language fluency checks. Furthermore, the speech model scoring module can continuously absorb new language data and user feedback through a continuous learning mechanism to improve its scoring accuracy and generation quality.

[0070] This application obtains corresponding multiple candidate text data by inferring the target speech features based on the large language model; then calls the preset speech model scoring module; and generates scores for each candidate text data based on the speech model scoring module; then filters out the designated candidate text data with the highest score from all the candidate text data; and subsequently uses the designated candidate text data as the target text. Based on the above processing flow, after inferring the target speech features based on the use of a large language model to obtain multiple candidate text data, this application will also intelligently evaluate the candidate text data based on the use of the speech model scoring module to filter out the required target text, effectively ensuring that the generated target text meets the requirements of natural language in terms of grammar, semantics and logic, improving the accuracy of the generated target text, and thus helping to improve user experience and interaction quality.

[0071] In some optional implementations of this embodiment, step S205 includes the following steps:

[0072] The target text is upsampled to obtain a corresponding processed text.

[0073] In this embodiment, the above-mentioned upsampling process includes: first pre-processing the target text line generated by the large language model to ensure that the text format meets the input requirements of the speech decoder (such as word segmentation, punctuation processing, etc.). Then, by increasing redundancy or adjusting the sampling rate, the text data is converted into a format that matches the input requirements of the speech decoder. For example, by inserting pause marks or adjusting the rhythm information of the text, the text data can be made more suitable for the subsequent speech generation process. This ensures that the upsampled text data is aligned with the processing rhythm of the speech decoder in the time dimension, avoiding speech generation errors or delays caused by data mismatch, and thus obtaining the corresponding processed text.

[0074] Based on a preset real-time processing strategy, the speech decoder is used to process the processing text to generate a corresponding speech feature sequence.

[0075] In this embodiment, the real-time processing strategy includes optimizing the computational efficiency and memory usage of the decoder to ensure real-time speech generation. For example, a lightweight neural network structure or model compression technology may be used to reduce computational latency in the decoder.

[0076] Specifically, the speech decoder can employ a streaming speech decoder (such as Tacotron 2 or FastSpeech) to receive upsampled text data and initialize its internal state in preparation for generating a speech feature sequence. Based on the aforementioned real-time processing strategy, the speech decoder employs a streaming approach, generating the corresponding speech feature sequence frame by frame. Each frame is generated based on the current text input and previously generated features, ensuring the continuity and naturalness of the speech.

[0077] The speech feature sequence is converted to obtain a corresponding speech waveform signal.

[0078] In this embodiment, the conversion process involves using a vocoder to receive the speech feature sequence generated by a streaming speech decoder, and then using a deep learning model to convert the abstract speech features into a playable speech waveform signal. This process involves detailed modeling of the speech signal, ensuring that the generated speech waveform is natural, smooth, and closely matches the original speech features.

[0079] The speech waveform signal is used as the reply speech.

[0080] In this embodiment, the generated reply voice can be temporarily stored in the system's memory or cache. This step ensures that the voice data can be quickly accessed during subsequent playback, reducing playback interruptions or freezes caused by data read delays. In addition, the cached reply voice can be converted to the necessary format (such as PCM encoding) to ensure compatibility with the playback device, thereby ensuring that the reply voice can be correctly played to the user.

[0081] This application obtains the corresponding processed text by upsampling the target text; then, based on a preset real-time processing strategy, the speech decoder is used to process the processed text to generate a corresponding speech feature sequence; the speech feature sequence is then converted to obtain a corresponding speech waveform signal; and the speech waveform signal is subsequently used as the reply speech. Based on the above processing flow, this application can automatically and accurately complete the decoding process of the target text through steps such as text upsampling, streaming speech decoding, and speech synthesis, so as to convert the generated text response into a natural and fluent reply speech, effectively ensuring that the system can generate high-quality speech responses in real time, thereby meeting the user's requirements for the naturalness and fluency of voice interaction.

[0082] In some optional implementations, step S206 includes the following steps:

[0083] Noise suppression is performed on the reply speech to obtain a corresponding first speech.

[0084] In this embodiment, the noise suppression process includes: first, analyzing the generated speech to identify the noise components therein. The noise may come from environmental noise, equipment interference, or imperfections in the generation process. Then, advanced noise suppression algorithms (such as spectral subtraction, Wiener filtering, deep learning noise suppression models, etc.) are used to suppress or eliminate the detected noise. These algorithms can distinguish between speech signals and noise and retain the main components of speech. The speech signal after noise suppression is subsequently evaluated to ensure that the noise level is significantly reduced while the main features of the speech signal (such as pitch and timbre) are retained.

[0085] Volume normalization is performed on the first speech to obtain a corresponding second speech.

[0086] In this embodiment, the volume normalization process involves first analyzing the volume level of the voice signal to identify portions that are too low or too high. Dynamic range compression technology is then applied to adjust the dynamic range of the voice signal's volume, making the overall volume level more uniform. This helps avoid auditory discomfort caused by sudden volume changes. Subsequently, a target volume level is set based on the application scenario and user needs, and gain adjustment or limiting techniques are used to adjust the voice signal's volume to the target range.

[0087] Perform clarity enhancement processing on the second speech to obtain a corresponding third speech.

[0088] In this embodiment, the clarity enhancement process includes improving the clarity of speech signals through spectral enhancement techniques (such as speech enhancement filters and harmonic enhancement). These techniques can enhance high-frequency components in speech, making it sharper. Furthermore, objective evaluation indicators (such as Short-Term Objective Intelligibility (STOI)) or subjective listening tests are used to evaluate the effectiveness of the speech clarity enhancement to ensure that the speech signal is easy to understand.

[0089] Performing a quality check on the third speech.

[0090] In this embodiment, the system monitors key quality indicators during the speech synthesis process, such as signal-to-noise ratio (SNR) and speech distortion, to determine whether these indicators are within acceptable ranges. If all indicators are detected to be within acceptable ranges, the third speech is determined to have passed the quality verification; otherwise, the third speech is determined to have failed the quality verification.

[0091] If the third speech passes the quality check, the third speech is used as the target reply speech.

[0092] In this embodiment, if abnormal speech quality is detected (such as severe noise or volume distortion), the system triggers an exception handling mechanism, such as regenerating the speech or adjusting optimization parameters. In addition, user evaluations of speech quality (such as clarity scores and satisfaction feedback) can be collected and integrated into the quality optimization process to continuously improve the parameters and strategies of algorithms such as noise suppression and volume normalization, thereby improving the overall quality of speech generation.

[0093] This application performs noise suppression processing on the reply voice to obtain the corresponding first voice; then performs volume normalization processing on the first voice to obtain the corresponding second voice; then performs clarity enhancement processing on the second voice to obtain the corresponding third voice; subsequently performs quality verification on the third voice; if the third voice passes the quality verification, the third voice is used as the target reply voice. Based on the above processing flow, this application performs noise suppression, volume normalization, clarity enhancement and quality verification processing on the reply voice, so as to achieve efficient and accurate quality optimization processing of the reply voice, thereby significantly improving the clarity and intelligibility of the generated target reply voice, and ensuring that the target reply voice can meet high-quality standards.

[0094] In some optional implementations, step S207 includes the following steps:

[0095] Get the device type of the voice device.

[0096] In this embodiment, the corresponding device type can be obtained by performing a type query on the above-mentioned voice device, for example, it may include a landline phone, a smart phone, a smart speaker, etc.

[0097] The target reply voice is adapted and adjusted based on the device type to obtain a corresponding feedback voice.

[0098] In this embodiment, the format and parameters of the target reply voice can be adjusted according to the device type of the voice device to ensure compatibility with the playback device, thereby obtaining the corresponding feedback voice.

[0099] Obtain a preset playback control strategy; wherein the playback control strategy includes a fluency control strategy and a synchronization control strategy.

[0100] In this embodiment, the fluency control strategy includes dynamically adjusting the voice signal transmission rate and playback cadence by real-time monitoring of network conditions and playback buffer status to avoid voice freezes or interruptions caused by network delays or jitter. The synchronization control strategy also includes ensuring that voice playback is synchronized with the user's interaction rhythm, for example, by immediately playing a response after a user asks a question or playing a prompt tone while the user is waiting, thereby improving the user experience.

[0101] Based on the playback control strategy, the feedback voice is transmitted to the voice device via a preset communication network.

[0102] In this embodiment, based on the aforementioned fluency control policy and the aforementioned synchronization control policy, the generated feedback voice can be transmitted back to the user's other voice device in real time via a communication network (e.g., PSTN, VoIP, etc.). During transmission, the voice signal can be encoded and compressed to optimize bandwidth usage.

[0103] The user receives the system-generated voice response through a voice device. The playback device converts the electrical signal into sound waves, and the system's response is received and understood through hearing, completing a complete interaction. If the user is satisfied with the response, the interaction ends; if the user has further questions, they can initiate a new interaction.

[0104] This application obtains the device type of the voice device; then adapts and adjusts the target reply voice based on the device type to obtain the corresponding feedback voice; then obtains a preset playback control strategy; wherein the playback control strategy includes a fluency control strategy and a synchronization control strategy; subsequently, based on the playback control strategy, the feedback voice is transmitted to the voice device through a preset communication network. Based on the above processing flow, after adapting and adjusting the target reply voice based on the device type of the voice device to obtain the feedback voice, this application uses the playback control strategy to utilize the communication network to play the generated feedback voice to the user in real time, which can effectively ensure the fluency and synchronization of the playback. In addition, this process ensures that the system can flexibly respond to the diverse needs of users and improve user satisfaction and interaction efficiency.

[0105] In some optional implementations of this embodiment, step S203 includes the following steps:

[0106] Call the preset voice adapter.

[0107] In this embodiment, there is no specific limitation on the selection of the above-mentioned voice adapter. A lightweight neural network module, such as a fully connected layer or an attention mechanism module, can be used according to actual business needs.

[0108] The voice feature is normalized based on the voice adapter to obtain a corresponding first voice feature.

[0109] In this embodiment, the normalization process may include mean-variance normalization.

[0110] Perform dimension matching processing on the first speech feature to obtain a corresponding second speech feature.

[0111] In this embodiment, the dimensionality matching process includes adjusting the feature dimensions to match the model input requirements of the large language model. By normalizing and dimensionality matching the speech features, it is possible to ensure that the feature representation is consistent with the data distribution during large language model training, thereby improving the generalization capability of the large language model.

[0112] The second speech feature is used as the target speech feature.

[0113] In this embodiment, the adaptively converted target speech features are passed to the input layer of a large language model (e.g., LLaMA-Omn i). This step can be achieved through an efficient feature transfer mechanism (e.g., memory mapping or direct data copy) to ensure that the feature data can be quickly and accurately received and processed by the large language model.

[0114] This application calls a preset voice adapter; then normalizes the voice features based on the voice adapter to obtain the corresponding first voice features; then performs dimension matching on the first voice features to obtain the corresponding second voice features; and subsequently uses the second voice features as the target voice features. Based on the above processing flow, this application performs normalization and dimension matching on the voice features based on the use of a voice adapter, thereby efficiently and accurately completing the adjustment of the voice features to convert the original voice instructions into feature representations suitable for large-scale language model processing, thereby ensuring the accuracy and standardization of the target voice features obtained. Lin Faxiang, this adjustment process not only retains the key information of the voice instructions, but also ensures the consistency of the feature representation with the model training data, laying a solid foundation for the subsequent text response generation.

[0115] In some optional implementations of this embodiment, after step S204, the electronic device may further perform the following steps:

[0116] Determine whether there is a historical dialogue text corresponding to the voice instruction.

[0117] In this embodiment, a historical conversation query of the user may be performed to detect whether there is a historical conversation text corresponding to the voice signal.

[0118] If so, call the default target cache.

[0119] In this embodiment, there is no specific limitation on the selection of the target cache, which can be determined according to actual storage requirements. For example, the system memory or cache can be used.

[0120] Based on the target cache, the target text and the historical conversation text are associated and stored.

[0121] In this embodiment, because interactions require contextual continuity (e.g., multi-round conversations), the system intelligently uses the target cache to associate the generated text response (target text) with the previously recorded conversation text. This allows the large language model to reference the complete conversation context during subsequent processing, thereby generating more accurate responses. Storing the target text ensures rapid access during subsequent speech decoder decoding, thereby reducing system response time increases due to data read latency.

[0122] This application determines whether there is a historical conversation text corresponding to the voice command; if so, calls a preset target cache; and subsequently associates and stores the target text with the historical conversation text based on the target cache. Based on the above processing flow, this application, through the use of a target cache, can automatically and intelligently associate and store the target text with related historical conversation text, improving the storage intelligence of the target text and enabling large-scale language models to reference the complete conversation context during subsequent processing, generating more accurate responses and thus improving the accuracy of response processing.

[0123] In some optional implementations, the user information obtained is obtained with the user's consent and complies with relevant laws and policies.

[0124] In addition, any software tools or components not provided by our company that appear in the embodiments of this application are merely examples and do not represent actual use.

[0125] In addition, this application is based on the LLaMA-Omni model architecture, which is designed for low-latency, high-quality voice interaction with a large language model, and the LLama-SpeakerPA model, which is retrained with insurance scenario data. The new model, LLama-Speaker PA, integrates a pre-trained speech encoder, a speech adapter, a large language model, and a streaming speech decoder. It eliminates the need for speech transcription and can simultaneously generate text and voice responses directly from voice commands with extremely low latency.

[0126] The description of innovation points includes:

[0127] 1. Multimodal Fusion Architecture: LLama-SpeakerPA achieves seamless speech-to-text and back-to-speech conversion by integrating a pre-trained speech encoder, speech adapter, large language model, and streaming speech decoder. This multimodal fusion architecture not only improves the system's interactive efficiency but also significantly reduces the system's response latency by eliminating the need for intermediate text transcription.

[0128] 2. Low-latency streaming voice decoding: Traditional voice interaction systems typically use serial processing, resulting in high latency. LLama-Speake rPA uses a streaming voice decoder, which generates corresponding voice responses in real time while the LLM generates text responses. This parallel processing mechanism significantly improves system response speed, reducing LLama-SpeakerPA's response latency to as low as 226ms, far lower than the 500ms or more of traditional systems.

[0129] 3. Targeted Dataset Construction: To better adapt to insurance business scenarios, this application has constructed a dataset containing 200,000 voice commands and their corresponding voice responses. This dataset was created by rewriting existing text command data and performing speech synthesis, ensuring the targeted and effective training of the model. Furthermore, based on different business scenarios, three major datasets can be distinguished, such as customer service, claims processing, and insurance consulting.

[0130] 4. Two-stage training strategy: LLama-SpeakerPA adopts an innovative two-stage training strategy. In the first stage, the training model generates text responses directly from voice commands; in the second stage, the model is further trained to generate voice responses. This staged training method not only improves the learning efficiency of the model, but also helps the model better adapt to specific business scenarios. Three specialized LLama-SpeakerPA models were obtained by training on the three major data sets mentioned in 3. The training process may include: 1) Stage 1: After the user inputs the voice, the content is converted into the large model through the voice encoder and adapter, and the adaptive layer and the large model inference layer are trained. 2) Stage 2: In the second stage, the adaptive layer and model layer are frozen, the generated data is upsampled, and then the reply content is generated through the decoder, and the voice is output through the Vocoder.

[0131] 5. Real-time Voice Interaction: LLama-Speaker PA enables real-time voice interaction in insurance services, including customer service, claims processing, and insurance consulting. Customers can communicate with the system in natural language over the phone, and the system accurately understands their voice commands and provides timely voice feedback.

[0132] 6. Adaptive Decoding Speed: LLama-SpeakerPA's streaming voice decoder can adjust the decoding speed based on actual needs, achieving different trade-offs between latency and voice quality. This adaptive mechanism enables the system to provide customized interactive experiences based on different business scenarios and customer preferences.

[0133] 7. Efficient Computing Resource Utilization: Compared to other speech-language models, LLama-SpeakerPA requires significantly less data and computing resources for training. Experimental results show that training LLama-SpeakerPA takes less than three days and requires four GPUs, significantly reducing the cost of model development and deployment.

[0134] In summary, this application provides an efficient and natural voice interaction solution for the insurance industry through its innovative multimodal fusion architecture, low-latency streaming speech decoding, targeted dataset construction, two-stage training strategy, and efficient computing resource utilization. This system not only significantly enhances the customer service experience but also improves business processing efficiency, bringing revolutionary changes to the insurance industry.

[0135] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0136] It should be emphasized that in order to further ensure the privacy and security of the above-mentioned target reply voice, the above-mentioned target reply voice can also be stored in a node of a blockchain.

[0137] The blockchain referred to in this application refers to a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (to prevent counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0138] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0139] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0140] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0141] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0142] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a speech processing device based on artificial intelligence, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0143] like Figure 3 As shown, the artificial intelligence-based speech processing device 300 described in this embodiment includes: a receiving module 301, an extraction module 302, an adjustment module 303, an inference module 304, a decoding module 305, an optimization module 306, and a transmission module 307. Among them:

[0144] The receiving module 301 is used to receive a voice command input by a user through a preset voice device;

[0145] An extraction module 302 is configured to extract features from the voice command based on a preset voice encoder to obtain corresponding voice features;

[0146] An adjustment module 303 is configured to adjust the voice features based on a preset voice adapter to obtain corresponding target voice features;

[0147] An inference module 304 is configured to perform inference processing on the target speech features based on a preset large language model to generate a corresponding target text;

[0148] The decoding module 305 is used to decode the target text based on a preset voice decoder to obtain a corresponding reply voice;

[0149] An optimization module 306 is configured to optimize the reply speech based on a preset quality optimization strategy to obtain a corresponding target reply speech;

[0150] The transmission module 307 is configured to transmit the target reply voice to the voice device based on a preset playback control strategy.

[0151] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0152] In some optional implementations of this embodiment, the reasoning module 304 includes:

[0153] An inference submodule, configured to perform inference processing on the target speech features based on the large language model to obtain corresponding multiple candidate text data;

[0154] The first calling submodule is used to call a preset speech model scoring module;

[0155] A generating submodule, configured to generate scores for each candidate text data based on the speech model scoring module;

[0156] A screening submodule, configured to screen out the designated candidate text data with the highest score from all the candidate text data;

[0157] The first determination submodule is configured to use the designated candidate text data as the target text.

[0158] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0159] In some optional implementations of this embodiment, the decoding module 305 includes:

[0160] An upsampling submodule, configured to perform upsampling processing on the target text to obtain a corresponding processed text;

[0161] A first processing submodule is configured to process the processed text using the speech decoder to generate a corresponding speech feature sequence based on a preset real-time processing strategy;

[0162] A conversion submodule, configured to convert the speech feature sequence to obtain a corresponding speech waveform signal;

[0163] The second determining submodule is configured to use the speech waveform signal as the reply speech.

[0164] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0165] In some optional implementations of this embodiment, the optimization module 306 includes:

[0166] A second processing submodule is configured to perform noise suppression processing on the reply speech to obtain a corresponding first speech;

[0167] A third processing submodule is configured to perform volume normalization processing on the first speech to obtain a corresponding second speech;

[0168] a fourth processing submodule, configured to perform clarity enhancement processing on the second speech to obtain a corresponding third speech;

[0169] A verification submodule, configured to perform quality verification on the third speech;

[0170] The third determining submodule is configured to use the third voice as the target reply voice if the third voice passes the quality check.

[0171] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0172] In some optional implementations of this embodiment, the transmission module 307 includes:

[0173] A first acquisition submodule is used to obtain the device type of the voice device;

[0174] An adjustment submodule, configured to adapt the target reply voice based on the device type to obtain a corresponding feedback voice;

[0175] The second acquisition submodule is used to acquire a preset playback control strategy; wherein the playback control strategy includes a fluency control strategy and a synchronization control strategy;

[0176] The transmission submodule is used to transmit the feedback voice to the voice device through a preset communication network based on the playback control strategy.

[0177] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0178] In some optional implementations of this embodiment, the adjustment module 303 includes:

[0179] The second calling submodule is used to call a preset voice adapter;

[0180] a fifth processing submodule, configured to perform normalization processing on the voice feature based on the voice adapter to obtain a corresponding first voice feature;

[0181] a sixth processing submodule, configured to perform dimension matching processing on the first speech feature to obtain a corresponding second speech feature;

[0182] The fourth determining submodule is configured to use the second speech feature as the target speech feature.

[0183] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0184] In some optional implementations of this embodiment, the artificial intelligence-based speech processing device further includes:

[0185] A judgment module, used to judge whether there is a historical dialogue text corresponding to the voice command;

[0186] A calling module, for calling a preset target cache if yes;

[0187] A storage module is used to associate and store the target text with the historical conversation text based on the target cache.

[0188] In this embodiment, the operations performed by the above modules or units correspond one-to-one to the steps of the artificial intelligence-based speech processing method in the aforementioned embodiment, and are not repeated here.

[0189] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0190] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0191] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0192] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the speech processing method based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0193] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions or process data stored in the memory 41, such as computer-readable instructions for executing the artificial intelligence-based speech processing method.

[0194] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0195] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0196] In an embodiment of the present application, first, a voice instruction input by a user through a preset voice device is received; then, based on a preset voice encoder, feature extraction is performed on the voice instruction to obtain corresponding voice features; then, based on a preset voice adapter, the voice features are adjusted and processed to obtain corresponding target voice features; and based on a preset large language model, the target voice features are inferred and processed to generate corresponding target text; subsequently, the target text is decoded and processed based on a preset voice decoder to obtain a corresponding reply voice; further, the reply voice is optimized based on a preset quality optimization strategy to obtain a corresponding target reply voice; finally, based on a preset playback control strategy, the target reply voice is transmitted to the voice device. Based on the above processing flow, the present application can communicate with users in real time and naturally by combining the use of a voice encoder, a voice adapter, a large language model, and a voice decoder, accurately understand the voice instructions input by the user, and provide high-quality voice feedback, effectively improving the processing efficiency of voice interaction, ensuring the quality of the generated target reply voice, and thus helping to improve customer service experience and business processing efficiency.

[0197] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned artificial intelligence-based speech processing method.

[0198] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0199] In an embodiment of the present application, first, a voice instruction input by a user through a preset voice device is received; then, based on a preset voice encoder, feature extraction is performed on the voice instruction to obtain corresponding voice features; then, based on a preset voice adapter, the voice features are adjusted and processed to obtain corresponding target voice features; and based on a preset large language model, the target voice features are inferred and processed to generate corresponding target text; subsequently, the target text is decoded and processed based on a preset voice decoder to obtain a corresponding reply voice; further, the reply voice is optimized based on a preset quality optimization strategy to obtain a corresponding target reply voice; finally, based on a preset playback control strategy, the target reply voice is transmitted to the voice device. Based on the above processing flow, the present application can communicate with users in real time and naturally by combining the use of a voice encoder, a voice adapter, a large language model, and a voice decoder, accurately understand the voice instructions input by the user, and provide high-quality voice feedback, effectively improving the processing efficiency of voice interaction, ensuring the quality of the generated target reply voice, and thus helping to improve customer service experience and business processing efficiency.

[0200] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0201] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A speech processing method based on artificial intelligence, characterized in that: The steps include: Receive voice commands input by the user through a preset voice device; Extracting features of the voice command based on a preset voice encoder to obtain corresponding voice features; Adjusting and processing the voice features based on a preset voice adapter to obtain corresponding target voice features; Performing inference processing on the target speech features based on a preset large language model to generate corresponding target text; Decoding the target text based on a preset voice decoder to obtain a corresponding reply voice; Optimizing the reply voice based on a preset quality optimization strategy to obtain a corresponding target reply voice; Based on a preset playback control strategy, the target reply voice is transmitted to the voice device.

2. The artificial intelligence-based speech processing method according to claim 1, characterized in that: The step of performing inference processing on the target speech features based on a preset large language model to generate a corresponding target text specifically includes: Performing inference processing on the target speech features based on the large language model to obtain corresponding multiple candidate text data; Call the preset speech model scoring module; Generating a score for each candidate text data based on the speech model scoring module; Filtering out the designated candidate text data with the highest score from all the candidate text data; The designated candidate text data is used as the target text.

3. The speech processing method based on artificial intelligence according to claim 1, characterized in that: The step of decoding the target text based on a preset voice decoder to obtain a corresponding reply voice specifically includes: Performing upsampling processing on the target text to obtain corresponding processed text; Based on a preset real-time processing strategy, the processed text is processed using the speech decoder to generate a corresponding speech feature sequence; Converting the speech feature sequence to obtain a corresponding speech waveform signal; The speech waveform signal is used as the reply speech.

4. The speech processing method based on artificial intelligence according to claim 1, characterized in that: The step of optimizing the reply voice based on a preset quality optimization strategy to obtain a corresponding target reply voice specifically includes: performing noise suppression processing on the reply speech to obtain a corresponding first speech; performing volume normalization processing on the first speech to obtain a corresponding second speech; performing clarity enhancement processing on the second speech to obtain a corresponding third speech; Performing a quality check on the third speech; If the third speech passes the quality check, the third speech is used as the target reply speech.

5. The artificial intelligence-based speech processing method according to claim 1, characterized in that: The step of transmitting the target reply voice to the voice device based on the preset playback control strategy specifically includes: Obtain the device type of the voice device; Adapting the target reply voice based on the device type to obtain a corresponding feedback voice; Obtaining a preset playback control strategy; wherein the playback control strategy includes a fluency control strategy and a synchronization control strategy; Based on the playback control strategy, the feedback voice is transmitted to the voice device via a preset communication network.

6. The artificial intelligence-based speech processing method according to claim 1, characterized in that: The step of adjusting the voice features based on the preset voice adapter to obtain corresponding target voice features specifically includes: Call the preset voice adapter; performing normalization processing on the voice feature based on the voice adapter to obtain a corresponding first voice feature; Performing dimension matching processing on the first speech feature to obtain a corresponding second speech feature; The second speech feature is used as the target speech feature.

7. The speech processing method based on artificial intelligence according to claim 1, characterized in that: After the step of performing inference processing on the target speech features based on the preset large language model to generate the corresponding target text, the method further includes: Determining whether there is a historical conversation text corresponding to the voice command; If so, call the preset target cache; Based on the target cache, the target text and the historical conversation text are associated and stored.

8. A speech processing device based on artificial intelligence, characterized in that: include: A receiving module, configured to receive voice commands input by a user through a preset voice device; An extraction module, configured to extract features from the voice command based on a preset voice encoder to obtain corresponding voice features; An adjustment module, configured to adjust the voice features based on a preset voice adapter to obtain corresponding target voice features; An inference module, configured to perform inference processing on the target speech features based on a preset large language model to generate a corresponding target text; A decoding module, configured to decode the target text based on a preset voice decoder to obtain a corresponding reply voice; An optimization module, configured to optimize the reply speech based on a preset quality optimization strategy to obtain a corresponding target reply speech; The transmission module is used to transmit the target reply voice to the voice device based on a preset playback control strategy.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the artificial intelligence-based speech processing method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based speech processing method according to any one of claims 1 to 7.