Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Voice source" patented technology

The voice source contains important lexical and non-lexical information. The non-lexical information can convey, for example, prosodic events, emotional status, as well as cues pertaining to the uniqueness of the speaker’s voice. In engineering applications, there is a need for a more accurate source model that could model different voice qualities.

Cross-scene forged voice detection method based on multi-level voice representation

PendingCN121884830ASpeech recognitionVoice sourceSupervised learning
The invention discloses a cross-scene forged voice detection method based on multi-level voice representation. According to the method, a voice sample is obtained from a to-be-detected voice source and input into a forged voice detection model subjected to end-to-end joint training for detection, and the model comprises a self-supervised learning voice representation module, a hierarchical time attention network and a lightweight classifier. The self-supervised learning speech representation module is used for extracting multi-level speech representations of speech samples, the hierarchical time attention network fuses the multi-level speech representations to obtain discriminant features, and the lightweight classifier outputs authenticity scores and forgery scores based on the discriminant features. And determining a detection conclusion of the voice sample according to the score result. Through hierarchical fusion and end-to-end joint training of multi-level speech representation, the stability and generalization ability of forged speech detection under a cross-scene condition are improved, and the detection accuracy and reliability are effectively improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Training method and system for self-supervised training predictor for speech separation

ActiveCN115762557Bimprove performanceSelf-supervised training feature improvementSpeech recognitionSound sourcesVoice source
The embodiment of the present application provides a training method and system of a self-supervised training predictor for speech separation. The method comprises: extracting self-supervised training features of each single person voice source speech by using a pre-training model; extracting shallow features for speech representation and deep features for context information in the self-supervised training features, and determining the shallow features and the deep features of each single person voice source speech as training labels of the self-supervised training predictor; inputting training mixed speech generated by each single person voice source speech into the self-supervised training predictor to obtain estimated features of each single person voice source speech; and training the self-supervised training predictor based on a loss function determined based on the estimated features and the training labels corresponding to each single person voice source speech. The self-supervised training predictor is trained and applied in a speech separation model in the embodiment of the present application, so that the accuracy of the self-supervised training features is improved, the performance of the speech separation system is improved, and the model parameters and the calculation complexity are reduced.
Owner:AISPEECH CO LTD

A multi-screen control sovereignty dynamic determination method based on in-vehicle voice source positioning

PendingCN122420709AVoice sourceSpeech sound
This invention relates to the field of dynamic control sovereignty determination technology, specifically to a multi-screen dynamic control sovereignty determination method based on in-vehicle voice source localization. The method includes acquiring the voice source localization result of the current in-vehicle voice command, generating the target semantic mapping path of the voice command, constructing control sovereignty candidates, and forming a control mapping matrix by combining the current screen status information. The method selects the screen node with the optimal response link as the current voice sovereignty screen, monitors the control behavior feedback information of the sovereignty screen node, and if no behavior feedback matching the voice command is detected within a preset time window, a sovereignty correction mechanism is triggered, and the sovereignty is re-determined based on the next-level preferred node in the mapping path. This invention improves the system's interactive fault tolerance capability under dynamic, variable, and abnormal conditions, and solves problems such as voice control link breakage, retry failure, and user experience interruption in existing technologies.
Owner:EAST EVER TECH CO LTD

Method for generating a spatial voice signal and a device thereof, method for receiving a spatial voice signal and a device thereof

A method for generating a spatial voice signal performed by a voice pickup device, the method comprises: obtaining, by a plurality of microphones associated with the voice pickup device, a plurality of audio signals; determining a number of voice sources based on the plurality of audio signals; extracting, for each of the voice sources, a voice signal based on the plurality of audio signals, so as to generate a plurality of voice signals corresponding to the number of voice sources; mapping the plurality of voice signals into a first synthesis signal corresponding to a first channel of the spatial voice signal and a second synthesis signal corresponding to a second channel of the spatial voice signal, wherein for each of the voice sources, a time delay exists between the first synthesis signal and the second synthesis signal, and the time delay is associated with a relative location between the voice source and the plurality of microphones; transforming the first synthesis signal and the second synthesis signal into a mono channel signal associated with the spatial voice signal.
Owner:HARMAN INT IND INC +1

Multilingual speech and semantic intelligent translation method and system applied to exhibition scene

The invention discloses a multilingual speech semantic intelligent translation method and system applied to an exhibition scene, and belongs to the technical field of machine translation, and the method comprises the following steps: S1, obtaining a multi-person question judgment result; s2, if the multi-person questioning judgment result is multi-person questioning, audio identification information is obtained through analysis, and otherwise, the audio identification information is directly obtained through analysis; s3, obtaining each storage question keyword, each contrast question keyword and a key matching weighting factor corresponding to each question keyword; s4, obtaining a same-group evaluation result, if the same-group evaluation result is the same group, analyzing to obtain the comprehensive matching similarity of the storage groups, otherwise, analyzing the comprehensive matching similarity of each parallel storage group; s5, obtaining a comprehensive matching judgment result, if the comprehensive matching judgment result is unqualified, performing secondary refining processing to obtain a question and answer, and otherwise, directly obtaining the question and answer; and S6, voice broadcasting is carried out, and accurate separation and language recognition of voice sources of different questioning users are achieved.
Owner:ZHEJIANG HUIZHAN ELF TECHNOLOGY CO LTD

A large model-based human-computer voice precise interaction method

The present application belongs to the technical field of voice interaction, and particularly relates to a human-machine voice precise interaction method based on a large model. In view of the problems of voice source confusion, low instruction recognition rate and insufficient interaction safety caused by the complex acoustic environment in the vehicle cabin, the present application fuses reverberation features and harmonic attenuation features to construct acoustic fingerprints, and realizes high-precision sound source positioning in combination with seat occupancy state verification; performs priority filtering and semantic pre-alignment on non-main driver voice, and eliminates fragmented interference; introduces a large language model to standardize the mapping of structured text into instructions conforming to syntax, format and semantic specifications, and embeds a dynamic permission verification mechanism based on sound source position. The scheme significantly improves the voice instruction recognition accuracy and system robustness, effectively suppresses false triggering, enhances user privacy protection, provides an efficient and precise voice interaction experience for the intelligent cabin, and has outstanding technical practical value and industrial application prospects.
Owner:JINAN BOSAI NETWORK TECH CO LTD

AI voice transcription input method and system based on engineering geology term model

The invention relates to the technical field of geological information, in particular to an AI voice transcription input method and system based on an engineering geological term model, and the method comprises the steps: processing a field environment through the combination of dual-microphone noise reduction and acoustic echo cancellation, and obtaining pure voice source data containing lithology and parameters; calling a terminology library and a model, translating the terminology library and the model into a tagged text, and correcting an easy-to-confuse expression; identifying association by means of a field association degree algorithm and a tree structure, and judging integrity; missing information is supplemented and recorded, and double input is supported through voice prompt and highlight pop-up window guidance; the data and the image are bound through a unique identifier, structured storage is carried out, and triple verification is carried out; the system comprises a data acquisition and transfer module, a dynamic interaction and prompt module, a multi-modal fusion and storage module and a data detection module which are respectively responsible for voice transfer, semantic completion, data storage and quality verification. According to the invention, precise transfer and automatic processing of field voice are realized, the term recognition accuracy is improved, the labor cost is reduced, and the data reliability is guaranteed.
Owner:POWERCHINA BEIJING ENG CORP

End-to-end speech separation algorithm based on speech language model

PendingCN121583282ASpeech analysisAudio restorationVoice source
The invention discloses an end-to-end speech separation algorithm based on a speech language model, and the algorithm comprises the steps: discretizing a continuous audio into a 32-order discrete codebook sequence through a residual vector quantization coder-decoder, and introducing a transcription start symbol lt through an SOT strategy; sOSgt, SOSgt; a special separator is lt; sCgt; and a termination symbol lt; eOSgt, EOSgt; splicing a multi-person voice sequence; extracting audio depth features by using a pre-trained WavLM model, and guiding an autoregression decoder to output a separated zero-order codebook sequence in combination with a cross attention mechanism; predicting a high-order codebook sequence step by step through a non-autoregression model, configuring an independent embedding layer to fuse low-order information, and introducing a task embedding mechanism to optimize modeling; based on a special separator lt; sCgt; and slicing the multi-order discrete codebook sequence, and outputting an independent voice source through an Encodec decoder. According to the method, the intelligibility of voice separation and the audio restoration quality can be effectively improved, the decoding speed is high, the subjective hearing experiment result and the downstream task performance are excellent, the scene that the number of speakers is unknown is supported, and the industrialization application prospect is wide.
Owner:SHANGHAI JIAOTONG UNIV

Speaker positioning method and device and storage medium

The invention discloses a speaker positioning method and device and a storage medium, and the method comprises the steps: collecting the micro-Doppler features of a target region, and extracting the throat vibration frequency information from the micro-Doppler features, so as to recognize whether an effective sound source exists in the target region or not; under the condition that the existence of the effective sound source is detected, collecting a depth point cloud for the target area, and at least extracting the three-dimensional coordinates of the mouth of the speaker by executing key point detection; whether synchronous consistency exists between the throat vibration frequency information and the mouth three-dimensional coordinates or not is verified; and when the synchronization consistency exists, determining the positioning information of the speaker according to the three-dimensional coordinates of the mouth. Therefore, the accuracy and the anti-interference capability of voice source positioning are improved by utilizing dual verification of physical vibration characteristics and spatial geometric information, and an efficient and stable speaker positioning effect can be kept in a complex and noisy environment.
Owner:AISPEECH CO LTD

Voice instruction recognition method and system for composite robot

The invention discloses a voice instruction recognition method and system for a composite robot, and the method comprises the steps: obtaining a plurality of voice signals, carrying out the preprocessing of the voice signals, obtaining independent voice instruction signals corresponding to different voice sources, extracting the voice features of the independent voice instruction signals, generating an instruction text, and carrying out the recognition of the voice instruction of the composite robot. And analyzing instruction semantics of the instruction text, matching the instruction semantics with pre-stored action task priority information, obtaining priority information of action tasks corresponding to the instruction semantics, respectively evaluating corresponding confidence coefficients, and calculating a multi-source priority information fusion score through a dynamic confidence coefficient weighted fusion mode. And selecting and executing an action task according to the multi-source priority information fusion score. According to the method, the literal meaning of an instruction can be understood, and the real urgency behind the instruction can be perceived, so that a decision conforming to an actual situation can be quickly and accurately made in an emergency, and the operation safety and the task continuity are ensured.
Owner:WUXI INSTITUTE OF TECHNOLOGY

Voice masking method, apparatus, and system, and vehicle

A voice masking method, apparatus, and system are provided, which are applicable to the field of intelligent vehicles. The method includes: determining a voice source location and a target masking location; receiving a sound signal from the voice source location, and detecting whether a voice signal exists in the sound signal, to generate a detection result; and if the detection result is that the voice signal exists in the sound signal, generating a masking sound for the voice signal, and outputting the masking sound to a first speaker at the voice source location and a second speaker at the target masking location; or if the detection result is that no voice signal exists in the sound signal, skipping generating the masking sound. According to this application, a requirement of an occupant for private voice communication in vehicle cockpit space can be met, and information security can be improved.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Human-machine voice precise interaction method based on large model

The invention belongs to the technical field of voice interaction, and particularly relates to a man-machine voice precise interaction method based on a large model. In order to solve the problems of voice source confusion, low instruction recognition rate, insufficient interaction safety and the like caused by a complex acoustic environment of a compartment, an acoustic fingerprint is constructed by fusing reverberation characteristics and harmonic attenuation characteristics, and high-precision sound source localization is realized by combining seat occupation state verification; performing priority filtering and semantic pre-alignment on the non-main driving voice, and eliminating fragmented interference; a large language model is introduced to standardize and map a structured text into an instruction conforming to grammar, format and semantic specifications, and a dynamic permission verification mechanism based on a sound source position is embedded. According to the scheme, the voice instruction recognition accuracy and the system robustness are remarkably improved, false triggering is effectively inhibited, user privacy protection is enhanced, efficient and accurate voice interaction experience is provided for the intelligent cockpit, and the technical practical value and the industrial application prospect are prominent.
Owner:JINAN BOSAI NETWORK TECH CO LTD

Artificial voice generation system

According to the present invention there is provided an artificial voice generation system designed to restore or augment speech for individuals experiencing voice loss or desiring an alternative vocal quality. The system comprises a controllable air pump, an airflow member delivering air to the user's oral cavity, and a sound generation member— such as a replaceable membrane cartridge—configured to vibrate and produce sound in response to airflow. A user interface allows manual or automatic modulation of airflow and voice parameters, enabling real-time adjustment of pitch, loudness, and voice quality, including male, female, or non-binary characteristics. The system may include sensors for pressure and airflow, anti-jamming features, and can be adapted for use with or without a neck stoma, making it suitable for a wide range of users, from laryngectomy patients to those with temporary or chronic voice loss. By providing a customisable, high-quality artificial voice source with improved naturalness and intelligibility, the invention addresses limitations of existing devices such as electrolarynxes and tracheoesophageal prostheses, offering a versatile and user-friendly solution for voice rehabilitation and enhancement.
Owner:LARONIX PTY LTD

An office worker emotion recognition method and system based on visible light and voice signals

PendingCN122451571AFacial movementVoice source
The application discloses an office staff emotion recognition method and system based on visible light and voice signals, relates to the technical field of computer vision and voice signal processing, and the method is characterized in that: face images and voice signals are synchronously collected, first, it is judged whether a staff is in a silent state or a speaking state, if the staff is in the silent state, emotion is recognized by analyzing facial motion unit features of the whole face, if the staff is in the speaking state, the voice source is confirmed by lip reading verification, and then emotion is recognized by fusing facial motion features based on only eyebrow and eye regions and voice acoustic features. The application effectively solves the interference problem of mouth movement on expression recognition when speaking, and improves the accuracy and robustness of emotion state monitoring.
Owner:HUNAN INST OF TECH

Data processing method and device, equipment, medium and product

PendingCN121865044ASelective content distributionEngineeringVoice source
The embodiment of the invention discloses a data processing method and device, equipment, a medium and a product, and can be applied to the technical field of data processing. The method comprises the steps of displaying a target voice comment in a comment display page, displaying a list access entry associated with the target voice comment in the comment display page in response to the fact that the voice content of the target voice comment belongs to a first content type, and displaying the list access entry associated with the target voice comment in the comment display page in response to a trigger operation for the list access entry. A first list page associated with the first content type is displayed, the ranking of interactive voices is displayed in the first list page, the voice content of the interactive voices belongs to the first content type, and the interactive voices originate from voices in the published content and / or voice comments of the published content. By adopting the embodiment of the invention, the interaction among multi-source voices can be improved.
Owner:XINGIN INFORMATION TECH (SHANGHAI) CO LTD