Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1140 results about "Speech processing" patented technology

Speech processing is the study of speech signals and the processing methods of signals. The signals are usually processed in a digital representation, so speech processing can be regarded as a special case of digital signal processing, applied to speech signals. Aspects of speech processing includes the acquisition, manipulation, storage, transfer and output of speech signals. The input is called speech recognition and the output is called speech synthesis.

Audio signal authenticity verification method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio signal authenticity verification method, device, equipment and medium, and the method comprises the steps: constructing an original audio text data set, and generating an adversarial sample set, inputting the original audio text data set and the adversarial sample set into an audio detection model for joint training to obtain an audio detection model subjected to adversarial training; the method comprises the steps of obtaining a to-be-detected audio signal and extracting an acoustic feature of the to-be-detected audio signal, obtaining a non-acoustic feature associated with the to-be-detected audio signal, constructing a multi-dimensional feature vector according to the acoustic feature and the non-acoustic feature, inputting the multi-dimensional feature vector into an audio detection model to generate an abnormal index, and executing a hierarchical response operation based on the abnormal index. According to the method, the robustness of the model is enhanced by introducing adversarial sample training, and the multi-dimensional feature vector is constructed by fusing the multi-modal features, so that accurate recognition and hierarchical response to the voice cloning attack are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video generation method and interaction method based on digital human, and device, storage medium and program product

Provided in the embodiments of the present application are a video generation method and interaction method based on a digital human, and a device, a storage medium and a program product. In the embodiments of the present application, text-to-speech processing is performed on the basis of voice features of a user and an emotion label, speech-to-expression processing is performed on the basis of a mapping relationship between the voice features of the user and expression coefficients, and a digital human model is rendered on the basis of speech signals and the expression coefficients, so as to obtain video data of the digital human model. Thus, voice features of a user are accurately simulated, so as to ensure that a speech output of a digital human sounds natural and is also highly personalized, thereby realizing personalized driving of the digital human, and improving the realism of the digital human in terms of voice and dynamic images. Thus, the user experience is improved, and the interactivity of the digital human and the authenticity and immersion are enhanced.
Owner:TAOBAO CHINA SOFTWARE

Speech feature processing method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice feature processing method, device, equipment and medium. Performing time resolution analysis based on the fused Mel band energy to generate a multi-scale Mel spectrum amplitude value, and performing nonlinear transformation on the multi-scale Mel spectrum amplitude value according to the noise intensity parameter to generate a noise suppression Mel component; and generating a perception weighting coefficient according to an auditory perception model, and executing frequency domain energy adjustment on the noise suppression Mel component to generate Mel spectrum representation. On the basis of frequency resolution self-adaption, time resolution dynamic adjustment and auditory perception modeling, nonlinear transformation and perception weighting processing are applied to the multi-scale Mel spectrum amplitude value, the influence of noise interference on voice features can be effectively reduced, and the key information retention capacity of voice signals is enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio coding and decoding method, device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an audio coding and decoding method, device, equipment and medium, and the method comprises the steps: carrying out the sliding window segmentation processing of an input audio signal, and generating signal segments; processing the signal segments through an encoder containing a multi-layer self-attention mechanism to generate continuous potential representations; performing decomposition vector quantization processing on the continuous potential representation to generate a discrete code; the discrete codes are processed through a decoder comprising a multi-layer self-attention mechanism, and reconstructed signal segments are generated; and splicing the reconstructed signal segments to generate a complete audio signal. According to the method, a traditional convolution structure is replaced by a multi-layer self-attention mechanism, global time sequence dependence modeling is carried out on the audio signals after sliding window segmentation, effective compression of potential representation is realized in combination with decomposition vector quantization, and continuity and fidelity of audio reconstruction are improved on the premise that calculation complexity is not increased.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice enhancement method and device based on noise perception, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice enhancement method, device and equipment based on noise perception and a medium. Environment feature information is extracted and input into an audio enhancement model to generate an enhanced audio signal; obtaining a reference audio sample, extracting a personalized feature vector, and carrying out personalized processing on the enhanced audio signal; and collecting playing feedback data, determining a playing time domain adjustment parameter and a playing frequency domain adjustment parameter, adjusting the personalized enhanced audio signal, and generating an optimized audio signal. According to the method, dynamic adjustment is realized in combination with the feedback parameters in the playing process by fusing the environmental perception information and the personalized speaker characteristics, clear and natural optimized audio output with personalized styles can be generated in a complex environment, and the voice interaction quality and adaptability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent (self-learning) subsystem in access networks

An intelligent (self-learning) sensor-aware and / or context-aware subsystem comprising (i) a System-on-a-Chip (SoC), (ii) a radio transceiver, (iii) a microphone, (iv) a voice processing module, (v) a (bio-inspired) neuromorphic event camera or a hyperspectral camera, (vi) a first set of computer implemental instructions in artificial neural networks (ANN) (which may include a transformer model / diffusion model or Poisson flow generative model++ (PFGM++)), (vii) a second set of computer implementable instructions to analyze and interpret contextual data and (viii) an autonomous artificial intelligence (AI) agent is disclosed.
Owner:MAZED MOHAMMAD A

Audio compression and reconstruction method, device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to service system platforms of medical health, financial science and technology, communication and the like, and discloses an audio compression and reconstruction method, device, equipment and medium. Residual vector quantization and inverse quantization are carried out to generate basic spectrum features; and compensating a high-frequency component by adopting a spectrum expansion network, predicting a time domain error in combination with a long-short-term memory network, generating a time domain residual compensation signal, and superposing the time domain residual compensation signal with the reconstructed time domain audio signal to obtain a target audio signal. Through residual vector quantization and neural network modeling, the audio reconstruction quality under low-bit-rate compression is improved, spectrum detail reservation and time domain error compensation are optimized, meanwhile, the calculation complexity is reduced, and the method is suitable for high-sampling-rate and resource-limited scenes.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice processing method and device based on voiceprint feature screening, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health, voice navigation and the like, and discloses a voice processing method, device, equipment and medium based on voiceprint feature screening. The method comprises the following steps: screening voice of a near-field main speaker based on voiceprint similarity and signal intensity, executing duration filtering and confidence verification, dynamically constructing a voiceprint feature library, generating a voiceprint mask matrix based on the voiceprint feature library, performing frequency band suppression on a to-be-processed voice signal, and outputting a purified voice signal. According to the method, the voiceprint library is dynamically constructed, the voiceprint mask matrix is generated based on the voiceprint feature library, frequency band suppression is performed on the to-be-processed voice signal, non-target voiceprints are effectively shielded, the voice signal quality is improved, and thus target voice is accurately captured in a high-noise environment.
Owner:平安科技(上海)有限公司

Knowledge distillation-based text-to-voice method, apparatus and device, and medium

The invention relates to the technical field of voice processing, can be applied to business scenes in the fields of medical health, financial science and technology, barrier-free service and the like, and discloses a text-to-voice method based on knowledge distillation, which comprises the following steps: carrying out standardization processing on an input text to generate a standard text sequence; the lightweight text encoder encodes the standard text sequence to generate a text implicit vector; the non-autoregressive acoustic feature prediction module maps the text implicit vector into a student acoustic feature sequence, and calculates alignment loss through knowledge distillation; and performing structured pruning and parameter quantization based on alignment loss, generating an acoustic feature sequence by the optimized model, and converting the acoustic feature sequence into a voice waveform by a vocoder. According to the method, through knowledge distillation, pruning optimization and parameter quantification, the reasoning speed and the cross-equipment adaptability are improved while the model size and the calculation requirement are reduced, so that the TTS system can realize efficient, low-delay and low-power-consumption voice generation in a resource-constrained environment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice emotion recognition method and device based on context information, equipment and medium

PendingCN120636474ASpeech recognitionSingle sentenceSpeech sound
The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a context information-based voice emotion recognition method, device, equipment and medium, which comprises the following steps: receiving an original voice stream and generating an independent voice segment, recognizing a text and determining a speaker role type, and extracting an acoustic feature index; and generating a preliminary emotion label, generating context information in combination with the historical dialogue text, and inputting the context information, the preliminary emotion label, the speaker role type and the acoustic feature index into a multi-modal fusion module to generate an emotion judgment result. According to the method, multi-modal fusion is realized on the basis of context information by combining voice, text and role information, so that the emotion change of each role can be accurately recognized and understood in a complex dialogue scene, the problems of large single sentence emotion judgment error and neglect of the context information in a traditional method are avoided, and the accuracy and stability of emotion recognition are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Beamforming using image data

A device capable of using image data for purposes of determining a location of a user and audio beam selection to isolate audio in the direction of the user. The beamforming / beam-steering may occur after determining the user's location in order to conserve computing resources that would otherwise have been spent determining beams for non-desired directions. The beamformed audio may be used for speech processing, a communication session involving the device, or other purposes.
Owner:AMAZON TECH INC

Speech enhancement method and device based on multi-scale feature learning, equipment and medium

The invention relates to the technical field of speech processing, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a speech enhancement method based on multi-scale feature learning, which comprises the following steps: framing an input audio signal, extracting a Mel-frequency spectrum feature, extracting a frequency domain feature by using a multi-scale convolutional neural network, and carrying out multi-scale feature learning on the frequency domain feature; carrying out coding dimension reduction on the image; noise is suppressed through a deep residual network, enhanced audio features are generated, a non-autoregression generative model is adopted for feature conversion, and finally a generative adversarial network is used for reconstructing a target voice waveform. Voice frequency domain features are extracted through the multi-scale convolutional neural network, and the feature expression ability of different frequency bands is improved; noise suppression is carried out through a deep residual network, and the purity of the voice signals is enhanced; feature conversion is optimized through a non-autoregression generation model, and the modeling efficiency of speech enhancement is improved; the target voice waveform is reconstructed through the generative adversarial network, and the naturalness and definition of the generated voice are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Environmental acoustic simulation voice generation method, apparatus and device, and medium

ActiveCN120612917ASpeech recognitionSpeech synthesisPhonetic environmentData set
The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses an environmental acoustics simulation voice generation method, device and equipment and a medium, and the method comprises the steps: obtaining and separating mixed voice data, and generating original voice content and original environmental acoustics information; converting the original voice content into first text information; determining a target environment acoustic tag in combination with the original environment acoustic tag, the first text information and the target geographical location information; acquiring target environmental acoustic information from a preset sound data set based on the tag, and adjusting the amplitude characteristic of the target environmental acoustic information to match the original environmental acoustic information; and synthesizing the adjusted target environment acoustic information and the original voice content into simulated voice data. According to the method, the target geographic position information is introduced to participate in acoustic feature determination and amplitude adjustment, so that the generated simulated voice data is more consistent in geographic semantics and acoustic performance, and the authenticity and concealment of voice environment disguise are effectively improved.
Owner:PING AN TECH (BEIJING) CO LTD

Detecting corrupted speech in voice-based computer interfaces

Approaches are generally described for corrupted speech detection in voice-based computer interfaces. First input data including first audio data representing a user utterance may be received. First data representing the first audio data may be generated using a first encoder. First text data representing a transcription of the user utterance may be generated. Second data representing the first text data may be generated using a second encoder different from the first encoder. Third data may be generated by combining the first data and the second data. The third data may be sent to a classifier network trained to predict a relevant corruption state for speech processing inputs. The classifier network may determine that the first input data corresponds to a first corruption state.
Owner:AMAZON TECH INC

System and method for enhancing speech of target speaker from audio signal in an ear-worn device using voice signatures

An ear-worn device is provided that operates to isolate and individually treat the received speech of a target speaker or multiple target speakers from an audio input signal detected in a multi-speaker environment. The ear-worn device uses a machine learning model that receives a voice signature of each of one or more target speakers as input signals, to identify and isolate the component of the audio input signal attributable to the target speaker(s). Once isolated, the target speaker's speech may be enhanced, de-emphasized, or otherwise processed in a manner desired by the wearer of the ear-worn device. The wearer may use an external electronic device, e.g., a phone, to select one or more target speakers in a conversation and / or configure various settings associated with processing the speech on the ear-worn device.
Owner:FORTELL RESEARCH INC

Voice style migration method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice style migration method, device, equipment and medium, comprising the following steps: acquiring a source voice signal and a target style feature, performing feature extraction on the source voice signal to generate a source voice content feature and a potential style feature, encoding the target style feature to generate an encoded target style feature, performing style decoupling and migration processing on the source voice content feature, the potential style feature and the encoded target style feature by using a pre-trained multi-modal large model, and generating a migrated feature, and generating a target voice signal based on the migrated features. According to the method, the semantic and style information of the source voice is fused, the encoded target style features are combined to execute style migration, effective decoupling and adaptive fusion of contents and styles are realized by using a multi-modal large model, and the cross-speaker and cross-scene voice migration effect and practicability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent conference content real-time translation method based on voice recognition

The invention relates to the technical field of voice processing, in particular to an intelligent conference content real-time translation method based on voice recognition, which comprises the following steps of: acquiring an original voice signal, generating an enhanced voice stream by adopting a sudden change noise suppression algorithm, and synchronously extracting a voiceprint fingerprint spectrum; inputting the enhanced voice stream into a hierarchical separation network, performing multi-speaker voice decoupling based on a voiceprint fingerprint spectrum, outputting a voice segment with an identity tag, and triggering updating of an incremental term knowledge base; generating a cross-language semantic consistency vector, and constructing a dynamically updated context memory pool at the same time; and converting the semantic consistency vector into a target language stream. According to the method, the spatial accuracy and semantic independence of voice decoupling are improved, a structured input basis is provided for subsequent semantic modeling and translation, and the method is particularly suitable for conference scenes with cross speaking and frequent overlapping speech streams.
Owner:广东公信智能会议股份有限公司

Hybrid speech processing method, electronic equipment and computer readable medium

The invention relates to the technical field of voice processing, in particular to a mixed voice processing method, electronic equipment and a computer readable medium. The method comprises the following steps: collecting mixed voice and environment influence parameters; carrying out bionic frequency domain analysis on the mixed voice to obtain low-frequency attenuation compensation characteristic data; performing multipath effect propagation analysis on the low-frequency attenuation compensation characteristic data through the environmental influence parameters to generate channel distortion data; performing time domain-frequency domain joint deconvolution processing on the mixed voice by using the channel distortion data to generate a direct sound component and a reflected sound component; performing adversarial training based on the direct sound component and the reflected sound component to generate anti-multipath speech enhancement data; and constructing a dynamic frequency compensation filter based on preset environmental acoustic characteristics. Through the multi-stage signal processing, frequency compensation and real-time optimization technology, the output quality of the mixed voice is improved.
Owner:GUANGZHOU ZHIYU CLOUD NETWORK COMMUNICATIONS CO LTD

Speech processing using user satisfaction data

Devices and techniques are generally described for generating user satisfaction data in a natural language processing system. In various examples, first user input data may be received by a natural language processing system. Behavioral data related to the first user input data may be determined. Natural language processing error data related to the first user input data may be determined. First response data corresponding to the first user input data may be determined. First response characteristic data related to the first response data may be determined. First user satisfaction data may be determined based at least in part on the first response characteristic data, the behavioral data, and the natural language processing error data.
Owner:AMAZON TECH INC

Speech semantic analysis method and device and storage medium

The invention discloses a speech semantic analysis method and device, and a storage medium. The method comprises the steps of determining a real-time speech recognition result corresponding to frame-by-frame input speech; inputting the real-time voice recognition result and the context information into a semantic integrity judgment model to judge whether the real-time voice recognition result forms a complete semantic unit or not; and determining a real-time semantic analysis result corresponding to the real-time voice recognition result based on the streaming semantic analysis engine and the context information under the condition that the real-time voice recognition result is detected to form a complete semantic unit. Therefore, by introducing a frame-by-frame real-time speech recognition mechanism and a semantic integrity discrimination model, a processing link of staged dependence and waiting in a traditional speech processing system is broken, and closer dynamic cooperation between speech recognition and semantic analysis is realized.
Owner:AISPEECH CO LTD

Voice text bidirectional conversion method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice text bidirectional conversion method, device, equipment and medium, and the method comprises the steps: respectively executing voice recognition or voice synthesis operation according to the type of input information; for the voice information, noise suppression parameters are generated in combination with the lip movement video data, noise reduction processing is executed, and the recognition accuracy is improved; for text information, a pre-generated speaker style vector is obtained, the vector is cited in the speech synthesis process to generate natural personalized speech, and lip movement information and tactile feedback which are synchronous with speech output are generated. According to the method, complex noise is suppressed by fusing lip movement data, personalized voice is generated by using the style vector, and lip movement and touch information is output, so that bidirectional real-time conversion of voice and text in a complex environment is realized, and recognition accuracy, voice naturalness and interaction synchronism are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-speaker dialogue voice analysis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses a multi-speaker dialogue voice analysis method, device and equipment and a medium, and the method comprises the steps: obtaining a to-be-analyzed multi-speaker dialogue voice, determining a naturalness score based on an acoustic feature, and obtaining a multi-speaker dialogue voice analysis result; determining a semantic consistency score based on voice embedding and semantic embedding corresponding to a preset text, determining a speaker consistency score based on embedding of a plurality of speakers of the same speaker, determining an interaction rationality score based on voice alternate overlapping duration, determining a diversity score based on a variance of voice features, and fusing the scores, the comprehensive mass fraction is obtained. According to the invention, through quantitative evaluation of five dimensions of naturalness, semantic consistency, speaker consistency, interaction rationality and diversity, a comprehensive quality scoring system is established, so that the evaluation result simultaneously reflects voice fluency, content matching degree, identity stability, interaction rhythm rationality and feature richness.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio data compression method and device, electronic equipment and storage medium

The invention discloses an audio data compression method and device, electronic equipment and a storage medium, relates to the technical field of voice processing, can be applied to financial science and technology and medical health business scenarios, and comprises the following steps: obtaining target audio data to be compressed; and inputting the target audio data into an audio coding and decoding quantization compression model to obtain a reconstructed audio signal corresponding to the target audio data, the audio coding and decoding quantization compression model being obtained through audio adversarial training. In the model, feature extraction and compression can be performed on target audio data by using a hierarchical neural network of an encoder to obtain a low-dimensional potential feature vector; a residual vector quantization module performs discrete quantization processing on the low-dimensional potential feature vector through a multi-layer cascaded codebook to obtain a discrete quantization code; and performing audio waveform reduction processing on the discrete quantization code by using a decoder to obtain a reconstructed audio signal corresponding to the target audio data. According to the invention, audio high-fidelity compression can be realized, and audio reconstruction tone quality is improved.
Owner:PING AN TECH (BEIJING) CO LTD

Conversation-based skill component for assessing a user's state

The present application provides techniques for implementing a skill component, configured to perform an assessment of a user, as part of a speech processing system. The system may receive a natural language user input requesting assistance. The skill component may, using one or more machine learning models, determine at least one characteristic of the natural language input (e.g., lexical embedding, acoustic embedding, topic, tone, etc.). The skill component may determine state data for a present session, where the state data indicates a topic of the natural language user input and / or a user state associated with the natural language user input. The skill component may determine past state data of one or more past sessions, and generate a question to the user based on the state data for the natural language user input and the past state data.
Owner:AMAZON TECH INC

Embedding-based large language model tuning

Systems and methods for embedding-based LLM tuning include generating a first embedding of received user input data and utilizing a translation model trained to associate user input with device names to generate a second embedding that differs at least in part from the first embedding. Reference embeddings corresponding to devices associated with user account data may be generated and a subset of the reference embeddings that satisfy a threshold similarity to the second embedding may be determined. A large language model (LLM) configured to determine a response to the user input data may utilize data for a subset of devices that correspond to the subset of the reference embeddings for performing speech processing.
Owner:AMAZON TECH INC

Data processing method and apparatus based on artificial intelligence, electronic device, computer program product, and computer readable storage medium

Embodiments of the present application relate to speech processing technology. The embodiments of the present application provide a data processing method and apparatus based on artificial intelligence, an electronic device, a computer program product, and a computer readable storage medium. The method comprises: acquiring speech, and acquiring emotion data corresponding to the speech; performing first speech content feature extraction processing on the speech to obtain a first speech content feature corresponding to the speech, and performing emotion feature extraction processing on the emotion data to obtain emotion features corresponding to the speech; performing fusion processing on the first speech content feature and the emotion features to obtain a fusion feature corresponding to the speech; and performing animation parameter mapping processing on the fusion feature to obtain controller parameters corresponding to the speech, wherein the controller parameters are used for controlling a virtual object model to be presented in a target image, and the target image matches the content of the speech and the emotion data.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice recognition processing method, system and equipment based on conference scene and medium

The invention relates to a voice recognition processing method, system and device based on a conference scene and a medium, and belongs to the technical field of voice processing. The voice recognition processing method comprises the following steps: acquiring an original conference audio stream collected by a microphone array; performing signal preprocessing on the original conference audio stream collected by the main channel, and outputting a pure voice signal; generating a sound source orientation thermodynamic diagram based on the original conference audio stream; extracting multi-dimensional voiceprint feature vectors from the pure voice signals, performing dynamic grouping, outputting a voice fragment set marked with voiceprint IDs, and generating an initial transcription text; dynamically correcting the initial transliteration text, and outputting a transliteration text stream with an industry term tag; and performing periodic memory enhancement processing on the transliteration text stream, outputting and analyzing a long text, and generating structured conference summary data. According to the invention, the automation level and accuracy of conference voice processing can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD

Two-process error correction method and device for real-time speech transcription

The invention provides a two-process error correction method and device for real-time speech transcription, and relates to the technical field of speech processing, and the method comprises the steps: extracting the Mel spectrum features of each segment, inputting each Mel spectrum feature into a lightweight end-to-end model, and obtaining a preliminary transcription text; splicing the segments according to a preset number to obtain a plurality of long segments, and inputting each long segment into a speech recognition model to obtain a high-precision transcription text; performing text comparison on the preliminary transcription text and the high-precision transcription text according to the confidence degree set of the preliminary transcription text to obtain all error vocabularies in the preliminary transcription text; and performing corresponding error correction processing on each error vocabulary in the preliminary transcription text according to the type of the error vocabulary and the high-precision transcription text to obtain a final transcription text. According to the method, through a two-process transcription error correction mechanism of the preliminary transcription text and the high-precision transcription text, transcription error accumulation is reduced on the premise that the real-time performance is not affected, and the transcription accuracy in a complex scene is improved.
Owner:NANJING DOLPHIN INTELLIGENT TECH CO LTD

5G communication equipment intelligent voice interaction method and system based on deep learning

The invention belongs to the technical field of voice interaction, and relates to a 5G communication equipment intelligent voice interaction method and system based on deep learning, and the method comprises the steps: generating a joint input signal containing voice frequency domain information and a motion state; outputting the de-noised voice segments and the corresponding environmental interference level parameters; generating a basic voiceprint ID containing a user biological feature identifier and a dynamic key fragment; calculating a cloud processing priority score according to the environmental interference level parameter and the equipment residual electric quantity value, when the score exceeds a preset threshold value, sending a voice processing request containing a basic voiceprint ID to an edge computing node, and otherwise, triggering a local semantic recognition process; and receiving a semantic recognition result from a cloud end or a local end, and generating a multi-round dialogue response instruction in combination with the context parameters in the historical interaction data of the user. According to the invention, the problem that local processing resource overload or cloud communication delay exceeding is easily caused by a fixed noise weight distribution mechanism is solved.
Owner:SHENZHEN BESNEL TECH CO LTD

Voice signal compression method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice signal compression method, device, equipment and medium, which comprises the following steps: executing Fourier transform on an initial voice signal to extract an amplitude spectrum and a phase spectrum, respectively processing the amplitude spectrum and the phase spectrum to generate features, and splicing the features to obtain a compressed voice signal; the method comprises the following steps: carrying out quantization by adopting a residual vectorization mode to generate a compressed feature vector, executing entropy coding to obtain a compressed code stream, recovering the compressed feature vector at a decoding side, reconstructing a spliced feature through residual reverse quantization, enhancing feature expression through up-sampling operation, and restoring a voice signal through inverse Fourier transform. According to the method, a dual-path processing structure of amplitude features and phase features is constructed, feature compression is realized in combination with residual vectorization and entropy coding, inverse quantization and up-sampling enhancement are executed in a decoding stage, a spectrum energy structure and phase continuity are effectively reserved, and reconstruction precision and fidelity of voice signals are improved under the condition of low bit rate.
Owner:PING AN TECH (SHENZHEN) CO LTD