Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Normal speech" patented technology

Normal speech between two people typically has a range of 50 to 60 decibels. When two people are speaking in a public place with background noise, normal speech is louder, around five extra decibels.

Voice interaction application test method and device and related equipment

The invention discloses a voice interaction application testing method and device and related equipment. The method comprises the steps of obtaining a target file stream from a preset file, wherein the preset file comprises a file stream generated after a normal voice interaction event stream of a voice interaction application is serialized; deserializing the target file stream to generate various voice type events; distributing the various voice type events to a voice interaction application; and driving the voice interaction application to execute the various voice type events to obtain a test result of the voice interaction application. According to the embodiment of the invention, the method does not involve the contact of an operation interface, is suitable for the test of a non-touchable voice interaction application, and guarantees the normal operation of the voice interaction application in the vehicle-mounted terminal.
Owner:BEIJING CO WHEELS TECH CO LTD

A deep learning-based teaching quality evaluation method and system

The present application belongs to the technical field of intelligent teaching, and particularly relates to a teaching quality evaluation method and system based on deep learning. The method comprises extracting feature data to be evaluated from normal speech, evaluating the feature data to be evaluated by using a speech evaluation model to generate an evaluation result, and determining the quality grade of the normal speech according to the evaluation result, wherein the feature extraction from noise speech and noise speech comprises extracting amplitude information and frequency information from a sound production section. The present application performs screening on teaching speech to identify abnormal sound sections; through voiceprint comparison, the teaching speech in which the abnormal sound sections that can match the pre-stored voiceprint are classified as noise speech, and the teaching speech that cannot be matched is classified as noise speech, so as to distinguish the noise speech originating from the background environment from the noise speech originating from the teaching subject, overcome the evaluation error problem caused by regarding the two as noise without distinction, and lay a data foundation for subsequent evaluation.
Owner:CNSCI SOFT EDUCATIONAL TECH (BEIJING) CORP

Interphone noise suppression method and noise suppression device

The invention provides an interphone noise suppression method and a noise suppression device, which are used for judging whether a sound signal is noise or normal voice by utilizing a decibel value of an actually acquired sound signal and / or a ratio of a voice signal to a total sound signal, and switching on or switching off an information channel between an interphone and an earphone according to a judgment result. Moreover, the decibel threshold value of the normal voice signal is dynamically updated according to the environmental data, so that the accuracy of noise discrimination and the noise suppression effect are improved. The noise suppression device designed by the invention is externally hung between the interphone and the earphone, when the noise suppression device is used for suppressing noise, the interior of the original interphone does not need to be disassembled or maintained, and the noise suppression device is particularly suitable for maintenance of old equipment interphones.
Owner:STATE OWNED SIDA MASCH MFG CO LTD

A false trigger suppression method for a speech recognition system

PendingCN122417020ACarrier signalAcoustics
This invention relates to the field of speech recognition technology, specifically to a method for suppressing false triggering in a speech recognition system. This method acquires and segments candidate wake-up speech signals, performs audible acoustic analysis and high-frequency carrier anomaly analysis on each speech segment to be verified, and calculates a command validity score by combining cross-segment consistency, audible injection risk, and target secondary discrimination threshold. Based on this, false triggering speech signals are suppressed. This invention combines candidate recognition, segment verification, audible acoustic analysis, and high-frequency carrier anomaly analysis, and comprehensively considers cross-segment physical consistency, audible injection risk, and target secondary discrimination threshold to generate a command validity score, thereby improving the accuracy, security, and reliability of false triggering recognition and normal speech command response.
Owner:BEIJING ZHONGWANG BOCAI TECHNOLOGY CO LTD

Phoneme recognition model training method, phoneme recognition method and phoneme recognition device

PendingCN121506109ASpeech recognitionPhoneme recognitionDysarthria
The invention relates to the technical field of speech recognition, and discloses a phoneme recognition model training method, a phoneme recognition method and a phoneme recognition device.A first sample phoneme sequence corresponding to normal speech data is recognized through a first phoneme recognition model, and a second sample phoneme sequence corresponding to dysarthria speech data is recognized through a second phoneme recognition model; and taking the second sample phoneme sequence as a reference phoneme sequence in a scene without a fixed text. According to the method, fine tuning training is performed on the basis of the first phoneme recognition model according to the individual dysarthria data of the dysarthria patient to obtain the individualized second phoneme recognition model, and the situation of free dialogue or poor text compliance can be covered even in a scene without a fixed text, so that the accuracy of recognition of the dysarthria is improved. The actual pronunciation ability of the patient can be reflected in a natural communication context, dependence on fixed texts and artificial phoneme level alignment labeling is avoided, higher flexibility and adaptability are achieved, and a reliable basis can be provided for rehabilitation training and curative effect tracking.
Owner:SHENZHEN UNIV +1

A speech recognition method based on air traffic control speech recognition model with speech rate perception

The present invention discloses a speech recognition method for an air traffic control speech recognition model based on speech rate perception, comprising: step 1: counting the normal speech rate range of air traffic control speech; step 2: constructing an air traffic control speech recognition model including the speech rate estimation layer; step 3: defining a joint loss function, training the air traffic control speech recognition model, and dynamically adjusting the CTC loss weight coefficient and the speech rate perception loss weight coefficient in the joint loss function during the training process; step 4: using the trained speech recognition model to recognize air traffic control speech data. The present invention introduces speech rate perception loss constraints to form a multi-task learning mechanism that combines the main task of speech recognition with the auxiliary task of speech rate regularity learning, which significantly improves the robustness and recognition accuracy of the air traffic control land-to-air call speech recognition model in complex environments, and provides more reliable technical support for the safe and efficient operation of air traffic control.
Owner:THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Spoken language evaluation method and device, equipment and storage medium

The invention relates to the technical field of audio processing, and discloses a spoken language evaluation method, the method is applied to spoken language evaluation equipment integrated with a speech speed prediction model, and the method comprises the following steps: obtaining a frame level acoustic feature of a to-be-detected audio, and calculating a speech speed value of the to-be-detected audio according to the frame level acoustic feature; inputting the frame-level acoustic features of the to-be-tested audio and the speech speed value of the to-be-tested audio into a speech speed prediction model to obtain a speech speed adjustment value of the to-be-tested audio; adjusting the speech speed of the to-be-tested audio according to the speech speed adjustment value of the to-be-tested audio; and scoring based on the to-be-tested audio after the speech speed is adjusted. After the speech speed of the audio to be detected is adjusted, the speech speed is more consistent with the normal speech speed, so that the start frame and the end frame of each audio clip corresponding to each phoneme of the reference text in the audio to be detected can be more accurately detected in the alignment process. Therefore, the alignment effect of the to-be-tested audio and the reference text is improved, and the scoring accuracy is improved. The invention further discloses a spoken language evaluation device and equipment and a storage medium.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Voice processing method and electronic device

Embodiments of the present application disclose a speech processing method and an electronic device, and relate to the technical field of information processing. The method comprises: obtaining original speech; preprocessing the original speech to obtain a plurality of original speech features; in the case where the plurality of original speech features comprise a first speech feature, extracting a voiceprint feature from the first speech feature to obtain a first voiceprint feature, the speech type corresponding to the first speech feature being a whisper speech type; determining a target voiceprint feature corresponding to the first voiceprint feature from a target voiceprint feature library; and performing whisper speech conversion on the first speech feature based on the target voiceprint feature to obtain converted normal speech. According to the present application, the same tone as the real tone of the user or the tone specified by the user can be restored in the whisper speech conversion process, thereby improving the user experience.
Owner:HONOR DEVICE CO LTD

Recording interference prevention method and device

The invention relates to the technical field of voice information security, and discloses a recording interference prevention method and device, and the method comprises an environment perception step, an equipment recognition and positioning step, an interference strategy decision-making step, a dynamic sequence generation step, a signal fusion construction step, a parametric array modulation step, a directional emission step, and an evaluation adjustment step. The device corresponds to the method. According to the application, different recording devices are accurately distinguished through an acoustic fingerprint identification technology, and a targeted interference strategy can be generated according to specific acoustic features of the devices; by combining an acoustic parametric array and a beam forming technology, accurate directional transmission of interference sound waves is realized, interference energy is enabled to act on target equipment in a concentrated manner, and the interference efficiency is improved; through fusion of multiple dynamic sequences and construction of composite baseband signals, the unpredictability and anti-filtering capability of interference signals are enhanced. Therefore, the high-efficiency interference performance is maintained, the influence on normal voice communication is reduced, and accurate and self-adaptive anti-recording protection is realized.
Owner:GUANGZHOU WO YUN ELECTRONIC PROD CO LTD

Voice modification detection using physical models of speech production

A computer may train a single-class machine learning using normal speech recordings. The machine learning model or any other model may estimate the normal range of parameters of a physical speech production model based on the normal speech recordings. For example, the computer may use a source-filter model of speech production, where voiced speech is represented by a pulse train and unvoiced speech by a random noise and a combination of the pulse train and the random noise is passed through an auto-regressive filter that emulates the human vocal tract. The computer leverages the fact that intentional modification of human voice introduces errors to source-filter model or any other physical model of speech production. The computer may identify anomalies in the physical model to generate a voice modification score for an audio signal. The voice modification score may indicate a degree of abnormality of human voice in the audio signal.
Owner:PINDROP SECURITY INC

An adversarial audio defense method and system based on audio conversion and transcription difference detection

The application discloses an anti-adversarial audio defense method and system based on audio conversion and transcription difference detection, and belongs to the technical field of artificial intelligence security. The method comprises the following steps: step 1, standardizing and multi-dimensionally converting an input original audio; step 2, calculating a character error rate between the original audio and an audio sample transcription result, and comparing the character error rate with a pre-set detection threshold; and step 3, triggering a security response mechanism when the audio contains potential adversarial interference, and submitting the audio to an automatic speech recognition model for transcription processing when the audio is normal speech. The application can realize rapid detection and interception of adversarial audio samples without modifying an existing automatic speech recognition model, significantly improves the stability and robustness of detection, enables the system to realize real-time detection with low delay, balances high accuracy and low false alarm rate, and meets the actual needs of a voice interaction scene.
Owner:NANJING UNIV OF SCI & TECH

Vehicle-mounted voice protection method, device and equipment and computer readable storage medium

The embodiment of the invention provides a vehicle-mounted voice protection method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring an in-vehicle voice signal in real time; analyzing the voice signal to obtain power spectrum distribution characteristics; according to the power spectrum distribution characteristics, multiple paths of periodic short pulse sequences are constructed; generating a multi-channel ultrasonic masking signal according to the multi-path periodic short pulse sequence; and transmitting the multi-channel ultrasonic masking signal by using an ultrasonic phased array, and forming an ultrasonic sound field in a target area in the vehicle. In this way, the directional interference characteristic of the ultrasonic phased array can be utilized, accurate noise injection of illegal recording equipment is achieved, and therefore vehicle-mounted voice privacy is effectively protected under the condition that normal voice communication is not affected.
Owner:CHERY AUTOMOBILE CO LTD

Audio processing method and system

The invention provides audio processing for assisting communication between a plurality of users of an online video game. The system is designed for when a first user is unable to use normal speech vi
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Noise reduction pickup equipment and noise reduction pickup methods

This invention provides a noise reduction and sound pickup device and method. The noise reduction and sound pickup device includes at least one noise pickup, a target sound pickup device, and a sound signal processing unit. The at least one noise pickup is positioned close to a corresponding noise source, suitable for acquiring ambient noise emitted by the noise source and generating an ambient noise frequency signal. The target sound pickup device is positioned close to a target sound source, suitable for acquiring the target sound from the target sound source and generating a target sound frequency signal. The sound signal processing unit inverts and amplifies at least a portion of the frequency signal in the ambient noise frequency signal, superimposes the inverted and amplified ambient noise signal with the target sound signal to obtain a useful sound, and outputs the useful sound. The noise reduction and sound pickup device and method can acquire and eliminate various complex noises emitted by large equipment, while simultaneously preserving and amplifying the required sound signal to ensure normal voice communication and improve communication quality.
Owner:CNNC ACCURAY (TIANJIN) MEDICAL TECH CO LTD

How to automatically switch between mesh calls and 5G data network calls

This application relates to a method for automatically switching between mesh calls and 5G data network calls. Specifically a Bluetooth device constructs a mesh network to form a mesh group One of the nodes is used as the master node, and an online loop is created on the AP P. All Bluetooth devices communicate with the application in a two-way dynamic heartbeat data mode, and the cloud synchronizes information in the virtual network formed by mapping the nodes; The voice data between online nodes is communicated in mesh mode; when the nodes of Bluetooth devices are offline or reconnected, the node device automatically switches to the data network mode through the application and performs network communication using the data network mode via the cloud. This automatic switching method between mesh calls and 5G data network calls monitors the online status of the device in real time and automatically switches the voice data transmission to the cloud network when the device / node goes offline, thereby ensuring the normal voice data transmission and reception of non - communicating nodes and ensuring the real - time nature, continuity and completeness of voice data transmission. ​
Owner:FUKA AIKOSHI INTELLIGENT TECHNOLOGY CO LTD

Information processing device, information processing method, computer program, learning device, remote conference system, and support device

Provided is an information processing device that perform processing related to speech conversion of a speech that is not normally uttered and does not include pitch information such as a whisper or a faint speech. The information processing device includes a speech-to-unit encoder that generates an acoustic unit from a speech waveform, and a unit-to-speech decoder that reconstructs a speech waveform from an acoustic unit. The unit-to-speech decoder is subjected to preliminary learning by self-supervised learning of a Masked Language Model type using a normal speech and a whisper without a text label of a specific speaker to generate an acoustic unit common to the normal speech and the whisper, the acoustic unit being a latent expression in which a difference between the normal speech and the whisper is absorbed.
Owner:SONY GROUP CORP

A voice wake-up interactive response method and system

ActiveCN119541487BSpeech recognitionBlind zoneTrigger Areas
A voice wake-up interactive response method and system are disclosed, primarily used to detect whether a user has spoken after a voice interaction system has been woken up. The given time window after wake-up is divided into different regions, and different techniques are used to process each region. Specifically, in the false trigger region immediately following the wake-up word detection, the system detects whether the user has spoken and determines whether the spoken words are confused with the ending sound of the wake-up word. In the normal speech detection region following the false trigger region, the system detects whether the user has spoken. In the blind zone detection region following the normal speech detection region, the system detects whether the user has spoken. If no spoken words are detected in the blind zone detection region, the system calculates the fundamental frequency. If the fundamental frequency calculation indicates the presence of spoken words for a certain duration, then spoken words are considered detected. Using the method and system of this invention can significantly reduce the proportion of false detections and false negatives.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

Pseudo-ear language generation method, device and equipment

The invention discloses a pseudo-ear language generation method, device and equipment. The pseudo-ear language generation method comprises the following steps: acquiring normal voice; extracting multi-scale acoustic features of the normal voice; based on the multi-scale acoustic features, determining sounding effort capable of representing sounding force and clearness; and performing joint modulation of the cross-domain acoustic parameters based on the sounding effort, wherein the joint modulation comprises at least two of a time domain, a frequency domain and an excitation domain. The acoustic parameters of the time domain are adjusted based on the sounding effort, so that dynamic fluctuation caused by semantic key points, emotion emphasis or breathing rhythm in real ear language can be reflected; the acoustic parameters of the frequency domain are adjusted based on the sounding effort, so that the phenomenon that the sounding is clearer by force can be simulated; the acoustic parameters of the excitation domain are adjusted based on the sounding effort, and the effect of enhancing the high-frequency hoarseness sound during forced blowing can be simulated; therefore, the naturalness and interpretability of the pseudo-ear language can be effectively improved.
Owner:SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

A speech detection method based on logarithmic graph Fourier transform feature extraction

ActiveCN119993192BSpeech analysisAlgorithmGraph fourier transform
This invention relates to the field of speech verification technology, and more particularly to a speech detection method based on logarithmic graph Fourier transform feature extraction, comprising the following steps: constructing a translation operator for the speech graph, using an exponential function to describe the decay of dependencies between speech samples, generating the Laplacian matrix of the graph, and representing the speech signal as an undirected graph to capture intra-frame and inter-frame structural relationships; mapping the sample values ​​of the speech signal to graph node signals, transforming the speech signal from the time domain to the graph frequency domain, extracting frequency domain features, and forming an enhanced feature representation by synchronously merging intra-frame and inter-frame oscillation analysis and combining time domain features; generating a detection score to determine whether the speech signal is a playback attack or belongs to normal speech. This invention, by introducing logarithmic graph Fourier transform and graph signal processing methods, effectively solves the limitations of existing technologies in playback speech detection, significantly improving the comprehensiveness of feature extraction, discriminative ability, and performance of the detection system.
Owner:NANJING UNIV OF POSTS & TELECOMM

An ultrasonic transducer drive circuit and electronic artificial larynx

The application discloses an ultrasonic transducer driving circuit and an electronic artificial larynx, and relates to the field of medical devices. In the ultrasonic transducer driving circuit, a power module is used for providing power supply for the whole driving circuit; a glottal wave pulse modulator is used for amplitude modulating a glottal wave signal of normal speech of a patient and a carrier signal for exciting the ultrasonic transducer to generate a modulated signal; a differential output power amplifier is used for signal conditioning and power amplification of the modulated signal, and the amplified driving signal is output in a differential mode; a transformer is used for boosting the amplified driving signal to generate a boosted driving signal; and a matching circuit is used for realizing impedance matching between the boosted driving signal and the ultrasonic transducer, so that the power of the boosted driving signal is maximized and loaded on the ultrasonic transducer to drive the ultrasonic transducer to emit ultrasonic waves. The ultrasonic transducer driving circuit can reduce the overall volume of the driving circuit and meet the portable requirement of a handheld electronic artificial larynx device.
Owner:BEIHANG UNIV

Alarm system

PendingJP2025149520AAlarmsEngineeringSpeech sound
To suppress waste and confusion attributable to the fact that first information whose output value has dropped is outputted after the elapse of a cessation period of a speech suspension mode.SOLUTION: An alarm system 1 includes an alarm device 2 and a server 3. The alarm device 2 can execute an ordinary speech mode and a speech suspension mode. In the ordinary speed mode, voice information transmitted from the server 3 is not subjected to storage processing but is outputted from a voice output unit 28. In the speech suspension mode, the ordinary speech mode is suspended over a predetermined period, and voice information is not outputted but stored through storage processing during the predetermined period. In the speech suspension mode, after the predetermined period has elapsed, additional output processing of outputting the voice information, which is stored through the storage processing, from the voice output unit 28 is executed, and resumption processing of the ordinary speech mode is executed. The alarm device 2 does not output first information, the real-time property of which is high, out of the voice information from the voice output unit 28, but outputs second information, the real-time property of which is low, from the voice output unit 28.SELECTED DRAWING: Figure 1
Owner:OSAKA GAS CO LTD

Audio conversion and transcription difference detection-based confrontation audio defense method and system

The invention discloses an adversarial audio defense method and system based on audio conversion and transcription difference detection, and belongs to the technical field of artificial intelligence security, and the method comprises the following steps: 1, carrying out the standardization processing and multi-dimensional feature conversion of an input original audio; step 2, calculating a character error rate between the original audio and the audio sample transcription result, and comparing the character error rate with a preset detection threshold; and step 3, when the audio contains potential adversarial disturbance, a safety response mechanism is triggered, and when the audio is normal voice, transcription processing is continued by the automatic voice recognition model. According to the method, on the premise that an existing automatic speech recognition model does not need to be modified, rapid detection and interception of the adversarial audio sample are achieved, and the stability and robustness of detection are remarkably improved; the system can realize real-time detection with low delay, high accuracy and low false alarm rate are both considered, and the actual demand of a voice interaction scene is met.
Owner:NANJING UNIV OF SCI & TECH

Nonlinear injection attack detection method and device based on hardware characteristics

The application discloses a kind of nonlinear injection attack detection method and device based on hardware characteristics, wherein, detection method includes the following steps: (1) the speech activity detection is carried out to the audio to be measured collected, and the audio to be measured is cut according to speech part, and a plurality of speech segments are obtained after eliminating non-speech part;(2) for each speech segment, simultaneously carry out undersampling audio detection and abnormal white noise detection;If there is undersampling audio similar to normal speech part and / or there is approximate white noise highly correlated with speech energy, it is determined that the speech segment is nonlinear injection, and a warning is issued to the user.The detection method in the application can be directly deployed on a smart device, and the detection device can be deployed near the smart device, which can independently complete the detection work and provide a convenient, universal and unavoidable nonlinear injection attack detection scheme for voice assistant users.
Owner:ZJU HANGZHOU GLOBAL SCI & TECH INNOVATION CENT

Composition and method for expediting dental anesthesia recovery

PendingUS20260021102A1Inorganic active ingredientsHydroxy compound active ingredientsDental anesthesiaDental surgery
A dental composition for expediting recovery from dental anesthesia comprises Vitamin C (75-2,000 mg), caffeine (30-400 mg), and beet root extract (250-7,000 mg) in synergistic combination. The composition addresses prolonged numbness following dental procedures by employing multiple complementary mechanisms: Vitamin C counteracts anesthetic effects through neurochemical pathways, beet root extract enhances vascular delivery via increased blood flow, and caffeine accelerates metabolic clearance. Administered as capsules immediately after or during dental procedures, the composition reduces post-procedural numbness duration, prevents soft tissue injuries from inadvertent biting, restores normal speech and eating function sooner, and promotes wound healing. The invention provides a safe, effective solution to a long-standing problem in dental anesthesia management.
Owner:DRESCHER DAVID